IRT adaptive testing reruns an agent benchmark at 38.5% of full cost with 1.03 percentage points of error
A production analytics agent serving tens of thousands of monthly active users was re-evaluated across 574 historical benchmark runs, split chronologically into calibration and held-out periods, comparing random sampling, historical caching, fixed representative subsets and item-response-theory adaptive testing. Multidimensional 2PL adaptive testing gave the best fidelity: 200 questions, 38.5% of a full run, produced 1.03 pp mean absolute error on the overall score. The team still shipped difficulty-stratified fixed subsets for operational simplicity, and showed those transfer to five other agent families without recalibration and stay stable on calibration windows as short as one day, which is the cheaper recipe for anyone rerunning an eval on every agent change.
↳ Follow the thread