Research
Only 22 of 113 DeepSWE Tasks Track Full-Benchmark Performance, So Most Agent Benchmark Runs Are Wasted Money
DeltaSelect (arXiv 2609.19607, 17 Sep 2026) resamples DeepSWE's published trials and finds only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.50 with full-benchmark performance, meaning most tasks tell you nothing about a candidate change. The open-source method picks tasks whose single-run result tracks the full benchmark, maps fractional verifier results to a common score by linear regression, and fits a fixed task set to a dollar budget. In a gpt-5.6-luna case study it drove skill and instruction revisions across 13 evaluations for $27.86 total, ending 58.1% cheaper per run ($1.75 versus $4.18, p=0.008) at a higher calibrated score (42.36% versus 36.46%).
↳ Follow the thread