Upgrade from LLM-as-judge to Agent-as-a-Judge for ~90% human alignment at a fraction of the cost
Agent-as-a-Judge: Evaluate Agents with Agents (OpenReview / ICML)·high signal
Agent-as-a-Judge gives the evaluator agentic tools (it can inspect intermediate steps, read artifacts, run checks) rather than scoring a final string, reaching ~90% agreement with human experts versus ~70% for conventional LLM-as-a-judge. On the DevAI benchmark it cut evaluation from ~86 hours / $1,297 to ~2 hours / $31 — about a 97% reduction. For agentic systems where the process matters as much as the output, the verifying agent catches failures a single-pass judge misses.