Sources
JEV-as-a-Judge paper: 277x cheaper than GPT-6 as an eval judge, and a confidence cascade keeps 99% of accuracy at 57% of the cost
arXiv 2609.26550 (submitted 2026-09-22, Li, Miao, Krishnan and Padman at CMU) compares Jev against sixteen generative and reward-model judges using blinded human adjudication. Jev costs $0.044 per 1,000 judgments at 152 ms median latency and lands within 3 points of the strongest LLM judge on RewardBench-style preference and HaluEval factuality. It trails by 14.5 points on JudgeBench, where the judge has to check a derivation or resist a well-written wrong answer. A frozen cascade that accepts confident verdicts and escalates the rest to GPT-6 Astra kept 99% of accuracy at about 57% of the fee, a pattern you can copy directly for eval pipelines.
Source
↳ Follow the thread