Back to articles

Tag archive

#failuremodes

T
Aug 20, 2026

The agent that trusted a bad API: silent failures and validation debt

An agent called an API, got a syntactically valid response, trusted it, and built three wrong decisions on top of it. No error was thrown. Every step succeeded locally. The cascade cost was 15× the original bad call — and all of it was preventable by a two-line validation check. This is validation debt: paying the cost of skipped checks in compounding failure downstream.

Aug 20, 20269 min read0 reactions0 comments
Y
Aug 19, 2026

Your timeout is a bet: pricing the tradeoff before you pick a number

Every per-step timeout is a number someone typed in without a model behind it — too short and you kill real work in flight, too long and you pay to sit idle waiting on a hang. Both mistakes are failure modes with a price tag. A small cost model finds the number that actually minimizes total cost instead of the one that felt safe.

Aug 19, 20266 min read0 reactions0 comments
T
Aug 19, 2026

The cost of undoing: partial failures and the cleanup bill

An agent writes to three systems, then fails on the fourth. The first three writes are now orphaned in a partially-succeeded state. Rolling them back costs more than the original operation — not in tokens, but in human time and coordination overhead. This is the cleanup bill: the hidden cost of partial failures that retry logic doesn't touch.

Aug 19, 20268 min read0 reactions0 comments
T
Aug 17, 2026

The cost of finding a failure after the customer finds it

An agent's failure isn't expensive because it failed—it's expensive because you found out about it from a customer complaint instead of an alert. Detection latency multiplies the cleanup cost by an order of magnitude. This post explores the arithmetic of failure-finding, why automated detection is worth the infrastructure cost, and how to choose detection strategies.

Aug 17, 20268 min read0 reactions0 comments
4
Aug 15, 2026

429 is not a timeout: why rate limits need their own retry budget

A 429 and a 500 both land in the same except block, so most retry budgets treat them the same: one bucket, one backoff curve, one circuit breaker. That conflation is wrong in both directions — it makes you wait too little for the failure that isn't yours, and panic too much over the one that is. Two error classes, two buckets, and the Retry-After header everyone reads and no one obeys.

Aug 15, 20265 min read0 reactions0 comments
P
Aug 14, 2026

Predicting agent failure before you ship it

A demo proves an agent can succeed once. It says almost nothing about how often it will fail under real load, real input distributions, and real adversarial garbage. The failures that cost you in production are predictable before release — but only if you test the things that actually shift between the demo and the deployment. Four pre-release signals that forecast production failure, and the ones that don't.

Aug 14, 20266 min read0 reactions0 comments
P
Aug 14, 2026

Postmortem: the agent that spent $200 retrying a 400

An agent burned ~$200 overnight retrying an HTTP 400 — a request that was defined to fail. No component was buggy; each layer retried "reasonably." The teardown: why retryability is a property of the error and not a default, how three nested retry caps multiply into 75 doomed attempts per item, and why per-step caps never bound a bill. With the two-line fix and a circuit breaker.

Aug 14, 20269 min read0 reactions0 comments
T
Aug 13, 2026

The caller gave up ten minutes ago: orphaned retries in agent fleets

A user closes the tab. An upstream request times out. A parent agent gets cancelled by its own budget. None of that reliably reaches the retry loop three calls deep, so the retry keeps going — burning tokens and rate-limit headroom for a result nobody will ever read. Cancellation is the one signal every fleet retry pattern assumes exists and almost none actually propagate.

Aug 13, 20265 min read0 reactions0 comments
Y
Aug 13, 2026

Your agent's failures are silent: measuring failure modes in production

Most agent failures don't throw. The run returns a result, exit code zero, and the result is wrong — or it burns an hour and quietly gives up. If your monitoring only counts exceptions, you're blind to the failures that actually cost you. A taxonomy of agent failure modes and the specific instrumentation that catches each one before your users or your bill do.

Aug 13, 20265 min read0 reactions0 comments
F
Aug 13, 2026

Failure modes in multi-agent teams: how a crew of agents breaks differently

A single agent fails by getting the task wrong. A team of agents fails in ways no single agent can: correlated collapse, diffused responsibility, context fragmentation, and consensus that converges on nothing. The four failure modes that only exist once you have more than one agent — and why adding agents can lower reliability.

Aug 13, 20265 min read0 reactions0 comments