JEV-as-a-judge cascade: accept confident verdicts, escalate the rest, and keep 99% of a frontier judge's accuracy
arXiv 2609.26550·medium signal
Compared against 16 generative and reward-model judges with blinded human adjudication, a decision-only Jev judge came within 3 points of the strongest LLM judge on preference and grounded factuality, at 0.36% of its fee. Its gap concentrated in low-confidence decisions and in derivation-checking or persuasive-wrong-answer cases. A frozen cascade that escalates only low-confidence items kept 99% of the comparator's accuracy. The pattern carries over to any judge: route by the cheap judge's confidence, not by task type.