Fetching from the wire…
Public story · 2026-08-25 · high
A three-stage check across 520 agent runs found language rewrites scoring 5.6, far below build-toolchain rewrites' 31.4.
Why now: SWE Refactor Bench's numbers are new to coverage as of August 25, with no separate release date given in the paper itself.
A new benchmark tested coding agents on 20 whole-repo stack migrations, and only 28 of 520 runs cleared every check. It's called SWE Refactor Bench. Detailed in a paper posted to arXiv, it grades each run through a migration audit, behavioral tests, and an independent verification agent.
Thirteen of the 20 tasks got no accepted solution from any model, and the best performer, claude-opus-5, scored 47 out of 100.
Build-toolchain rewrites scored 31.4 out of 100. Language rewrites scored lower, at 5.6.
The paper names the dominant failure mode Blindness. Of the 340 runs that cleared the audit stage, 58% reached 99% of fixed checks, but only 26% reached 100%. That last point is where most runs die, invisible to a test count that only shows how close a run looks, not whether it's done.
I've seen this pattern with my own agent use. A refactor's tests go green, and then a second pass catches something the tests missed. This benchmark puts a number on how often that gap causes a failed run.
Each link below shares sources, entities, or timing with this story.
Shared entity: Best / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Best; reported by the same outlet (arxiv.org); overlapping topics (agent, best, test).
Shared entity: Build / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Build; reported by the same outlet (arxiv.org); overlapping topics (agent, check, runs).
Shared entity: Best / Same source domain / Shared topic / Tension
Both cover Best; reported by the same outlet (arxiv.org); overlapping topics (agent, best, check).
Shared entity: Best / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Best; reported by the same outlet (arxiv.org); overlapping topics (agent, best).
Shared entity: Best / Same source domain / Shared topic / Earlier coverage / Downstream implication
Both cover Best; reported by the same outlet (arxiv.org); overlapping topics (agent, best).
Shared entity: Best / Same source domain / Shared topic / Earlier coverage
Both cover Best; reported by the same outlet (arxiv.org); overlapping topics (accepted, audit, best).
Both cover Best; reported by the same outlet (arxiv.org); overlapping topics (agent, best, test).
Shared entity: Build / Same source domain / Shared topic / Earlier coverage
Both cover Build; reported by the same outlet (arxiv.org); overlapping topics (agent, only).