Back to articles

Tag archive

#llmasjudge

W
Aug 22, 2026

What to actually measure when your agent "works"

"It works" is a demo result, not a measurement. An agent is a trajectory, not a function, and grading only the final answer throws away most of what decides whether it's reliable. Here's a five-layer scheme for what to measure — outcome, trajectory, cost, failure class, and stability under nondeterminism — with a small harness that computes it.

Aug 22, 20269 min read0 reactions0 comments
Y
Aug 12, 2026

Your LLM-as-judge is lying to you

A model grading your agent's output is the only thing that scales for subjective quality — and it's a biased instrument you're reading as a ruler. Here are the biases that actually move scores (position, verbosity, self-preference, leniency), why raw agreement with a human hides them, and how to validate and harden a judge with code — including why you should be reporting Cohen's kappa, not accuracy.

Aug 12, 20268 min read0 reactions0 comments
E
Aug 10, 2026

Evaluating your evals: how to know the LLM judge is right

You built an LLM judge to grade your agent. What grades the judge? An eval you never validated is a ruler you never checked against a meter stick — and a biased judge doesn't just add noise, it moves your headline number in a consistent direction. How to meta-evaluate a judge: the labeled set, the agreement metric that isn't accuracy, and the drift check.

Aug 10, 20265 min read0 reactions0 comments
LLM-as-judge — 모델이 모델을 평가할 때 무엇이 깨지고 무엇이 살아남는가
Jun 6, 2026

LLM-as-judge — 모델이 모델을 평가할 때 무엇이 깨지고 무엇이 살아남는가

광고 카피 자동 생성·RAG 답변 품질·챗봇 응답 평가는 사람이 다 못 봅니다. LLM에게 "이 출력이 좋은가"를 물어 점수를 받는 LLM-as-judge가 표준이 되어가지만, 그 자체가 깨지는 자리도 많습니다. position bias·verbosity bias를 알고 보정하는 운영법.

Jun 6, 20262 min read0 reactions0 comments