Item Response Theory Fit to 8 Safety Benchmarks Across 192 Models Cuts Evaluation Cost 97-99% and Detects Naive Sandbagging
In what the authors call the largest psychometric analysis of LLM safety evaluations to date, IRT models fit to eight safety benchmarks across 192 language models find that three interpretable factors — refusal strictness, truthfulness, and contextual harm — explain most between-model variance. Psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, with roughly ten adaptively chosen items sufficing for several individual benchmarks, a 97-99% cut in evaluation cost. IRT also supports per-model audits, detecting naive sandbagging and silent model swaps behind an API, which is directly usable by anyone maintaining an eval harness against a hosted endpoint.
↳ Follow the thread