Research
Mechanically Truncating Old Tool Results Lifts SWE-bench Verified Fail-to-Pass From 28% to 49% on Unchanged Weights
arXiv 2608.26218 holds the model and task fixed and changes only the harness, comparing full time-ordered conversation against a variant that shortens older tool results as context fills and reacts to repeated or stalled work. On 169 SWE-bench Verified tasks with a 20,480-token window and a fixed 480-second endpoint, mean per-task fail-to-pass fraction goes from 28% to 49% and complete solutions from 43 to 72, and the same frozen treatment transfers to three other model designs without retuning. The authors argue coding-agent evaluations must name the model and harness together as the tested solver.
↳ Follow the thread