
How Arbiter works, why it was built, and the four measurements it made about an LLM judge that I had to look at twice.
Your eval harness trusts a witness it never cross-examined Every agent benchmark I have...
Tag archive

Your eval harness trusts a witness it never cross-examined Every agent benchmark I have...

Rewind: counterfactual replay for agent pipelines Every agent trace tool I've used has the...

A teardown of **Checkpoint: a 30-step agent run on Temporal that survives kill -9, with an...

Most agent memory today is one of two failure modes. Either it's an append-only log that grows until...

Your agent's CLAUDE.md is lying to it, and nothing is checking Every team I talk to has...

How I built a single React client that renders a Mastra agent and a hand-rolled emitter with zero...

A technical walkthrough of a generative-art model-arena: sandboxed p5.js iframes, honest pixel...

Ask two models the same open question and you get two walls of prose that are nearly impossible to...

A technical deep-dive into Cascade: a two-model arena where the benchmark is a constraint — pure...

How to build an app that runs AI-generated music code in sandboxed iframes, measures it with a real...

Model benchmarks are numbers on a leaderboard. Nobody feels them. I wanted the opposite: send one...

Liquid syntax error: Unknown tag 'endraw'