An Agent Benchmark Compressed to 20 Tasks Predicts Full Scores With 14-28% Less Error Than the Best Baselines
arXiv 2609.18909 (16 Sep 2026) tackles the cost of agent evaluation, where existing benchmark-compression methods model redundancy only in task-model final-score distributions. DualViewEval analyzes large-scale trajectories, identifies six complementary process signals associated with final performance, and jointly exploits outcome and process relations to learn an exact-size miniset that predicts full-benchmark scores. Across five agent benchmarks and five baselines it wins on all datasets, hitting 24x to 40x compression on APEX-Agents and BFCL with only 20 tasks, cutting mean absolute error 14.5-28.2% over the strongest competitors and improving Kendall's tau by up to 7.2% relative to EssenceBench on SWE-bench Verified.
↳ Follow the thread