Sources
Ten LLM judges carry only as much statistical information as 3.5 independent ones
This paper measures the error correlation that consensus-among-judges evaluation assumes away, finding an average pairwise error correlation of 0.21 across a bank of ten open-weight and frontier judges, which collapses their effective sample size to roughly 3.5 independent judges. The dependency is strongest among high-accuracy frontier judges, including ones from different providers, which kills the usual mitigation of mixing vendors. In up to 28% of comparisons, ignoring shared errors produces a conclusion that one system is significantly better when accounting for them does not, so anyone running an LLM-judge eval gate is likely overstating their confidence intervals.
Source
↳ Follow the thread