We spent thirteen weeks about to buy a bigger database
Our service dashboard had a panel called database time, and for thirteen weeks it showed a p99 of...
Tag archive
Our service dashboard had a panel called database time, and for thirteen weeks it showed a p99 of...
We shipped JSON logs and called it observability. Then checkout broke and nobody could say if it was us or the payment provider, or since when.
GitLab's half-year Co-Create program recap flags three shipped CI/CD changes: new W3C Trace Context variables in pipeline jobs, configurable merge-train pipeline limits, and a REST API for Terraform state protection rules.
An agent loop that only prints iteration=7 is already unbounded. The process can die, two retries can...
You instrumented the services. A request comes in at the edge, crosses four of them, hits the...
We had spent a quarter rolling out distributed tracing. Context propagation through eleven services,...
A CNCF write-up walks through wiring the OpenTelemetry Collector's alpha githubreceiver to org-level Actions webhooks, turning workflow_run and workflow_job events straight into OTLP spans.
A customer sent us a screenshot of a 500 error with a timestamp. That was all we had. The request had...
LLM observability: why traces, not just logs You probably already log responses from your...

The Last Piece of the Puzzle The previous four articles dissected the agent's control...

Using Langfuse to trace multi-step agent workflows, replace custom eval logic, and consolidate LLM observability into one tool that actually scales.

Distributed Tracing: Following a Request Across Microservices A practical guide to...