Research
Covering Every Task Beats Uniform Resampling: 87% Lower Score MSE on a 16-Model LiveCodeBench Replay
arXiv 2609.29140 derives tight bounds on how much replication a fixed-budget repeated eval needs to certify a narrow confidence interval. In an equal-budget LiveCodeBench replay with 16 models, 880 tasks and five outputs per task, a design that covers every task cut median point-estimate MSE by 87.0% against pooled uniform sampling. Its joint mean and disagreement certificate gave narrower intervals in 15 of 16 panels, with median width down 30.6%. The advice for anyone running pass@k evals is to spend budget on task coverage before resampling.
Source
↳ Follow the thread