Agent-as-a-Judge: Evaluate Agents With Agents
How to build agent-as-a-judge evaluation into production systems, with patterns for scoring, guardrails, cost control, and failure modes.
Tag archive
How to build agent-as-a-judge evaluation into production systems, with patterns for scoring, guardrails, cost control, and failure modes.
"It works" is a demo result, not a measurement. An agent is a trajectory, not a function, and grading only the final answer throws away most of what decides whether it's reliable. Here's a five-layer scheme for what to measure — outcome, trajectory, cost, failure class, and stability under nondeterminism — with a small harness that computes it.
Why verifying coding-agent output is now harder than generating it: reward hacking, the verification 'impossible triangle', and four verifier families that beat it.
How to evaluate AI agents in 2026: grade the trajectory not the answer, build regression sets from real failures, and report pass^k reliability.
A model grading your agent's output is the only thing that scales for subjective quality — and it's a biased instrument you're reading as a ruler. Here are the biases that actually move scores (position, verbosity, self-preference, leniency), why raw agreement with a human hides them, and how to validate and harden a judge with code — including why you should be reporting Cohen's kappa, not accuracy.
You built an LLM judge to grade your agent. What grades the judge? An eval you never validated is a ruler you never checked against a meter stick — and a biased judge doesn't just add noise, it moves your headline number in a consistent direction. How to meta-evaluate a judge: the labeled set, the agreement metric that isn't accuracy, and the drift check.
Using one AI to grade another is now common — but the biggest audit yet shows these graders are consistent without being correct. A judge that always picks "answer A" scores perfectly on consistency.
A practical 2026 guide to evaluating RAG: retrieval vs generation metrics, the exact RAGAS classes, where LLM-as-judge lies to you, and a runnable CI gate.
An un-seeded LLM judge is a coin flip even at temperature 0. How we made signed eval verdicts replayable instead of attesting noise.

광고 카피 자동 생성·RAG 답변 품질·챗봇 응답 평가는 사람이 다 못 봅니다. LLM에게 "이 출력이 좋은가"를 물어 점수를 받는 LLM-as-judge가 표준이 되어가지만, 그 자체가 깨지는 자리도 많습니다. position bias·verbosity bias를 알고 보정하는 운영법.
The Illusion of Precision When a benchmark report declares that Model A scores 87.3% on...
Traditional RAG evaluation relies on human-annotated "standard answers," but in the GraphRAG era,...