Skills
Coding agents pass whole-repo migration tests by copying the original code, and only 5.4% of runs survive a three-stage check
SWE Refactor Bench covers 20 whole-repository stack migrations across 4 technical-debt categories and grades each run through a Migration Audit, behavioural tests, and an independent verification agent. Only 28 of 520 runs (5.4%) pass all three stages, 13 of the 20 tasks got no accepted solution at all, and the best model, claude-opus-5, scores 47.0 out of 100. The named failure mode is Blindness: of 340 runs that clear the audit, 58% reach 99% of fixed checks but only 26% reach 100%, and agents do far better on build toolchain rewrites (31.4) than language rewrites (5.6).
↳ Follow the thread