Every LLM Judge Grades Stronger Models More Leniently, and a Label-Free Ensemble Tracks an Oracle Within 0.5 Points
Across four benchmarks and 36 judge-examinee pairs in absolute-scoring settings, a model's own task accuracy predicts its judging accuracy at Pearson r of 0.90 or better and inversely predicts directional bias at r below -0.83, but accuracy does not buy fairness: more capable examinees get more lenient judgments from every judge at r of 0.83 or better. Calibrated weighted majority voting estimates each judge's false-positive and false-negative rates purely from inter-judge disagreement, needing no ground-truth labels or task metadata. Under shifting task distributions it stays within 0.5 percentage points of an oracle with perfect error-rate knowledge, beating both single judges and plain majority voting.
↳ Follow the thread