Skills
LLM judges of tool-calling all collapse into a 77-82% band on hard tasks, and giving them the answer key makes strong models worse
arXiv 2608.26623 (EMNLP 2026) built AgentJudgeBench from 3,808 instances stratified by difficulty and found judge alignment degrades monotonically with task difficulty, 1.5x faster when no ground truth is available. On hard queries without ground truth, all six judges converged into a 77-82% band, which points at a workflow-complexity ceiling rather than a model-capacity one. Supplying reference answers degraded the stronger models through over-anchoring; structured rubrics were the only intervention that helped, at up to 6.5 points, while chain-of-thought and temperature changes did nothing.
↳ Follow the thread