Research
Agent Benchmarks Carry a Double Measurement Confound: Moving Execution Decisions Off the Scaffold Turns a Flat Leaderboard Into a Spectrum
arXiv 2609.09218 (submitted 6 Sep 2026) names two ways an agent benchmark score can fail to measure the model: execution-critical decisions get made by a fixed scaffold rather than the model, and the scorer grades shape rather than task correctness. The audit-and-repair protocol transfers execution decisions to the model, swaps shape-based scoring for seeded ground truth, and reports worst-case and tail-risk metrics instead of just the mean. On ComtradeBench the joint intervention turns a nearly flat leaderboard into a reliability spectrum, and auditing existing benchmarks shows scorer validity is benchmark-specific while scaffold ownership was an uncontrolled axis everywhere they probed.
↳ Follow the thread