Skills
Renaming identifiers and adding dead code costs code agents up to 6.7 points of resolve rate, and the simpler scaffold holds up better
Semantics-preserving transformations of SWE-bench repos, control-flow rewrites, dead-code injection, and identifier renaming, produced statistically significant degradation in 6 of 16 model-scaffold-dataset configurations, up to 6.7 percentage points mean resolve-rate drop. Rankings flipped by scaffold: Qwen looked robust under mini-SWE agent and brittle under OpenCode. The mini-SWE agent, the simpler scaffold, was more robust overall, which argues against assuming a heavier harness buys reliability.
↳ Follow the thread