Dispatch
Allen AI's BenchMIRT finds you can throw away 90% of a benchmark's questions and keep the same model rankings
BenchMIRT applies multidimensional item response theory to 100 open-weight models scored on 16 benchmarks totaling 34,000+ questions, without being told which benchmark measures what. Keeping just 10% of questions generally preserved model rankings and 50% often matched the full benchmark, while held-out question prediction hit 79% accuracy against a 70% baseline. It also surfaced mislabeled evals: BBQ (social bias) tracks general reasoning more than safety, and WMDP tracks reasoning inversely, meaning higher-reasoning models score lower.
↳ Follow the thread