Skills
254 SWE-bench submissions audited: the top two entries both resolve 396/500, and no adjacent top-thirty pair on Verified is statistically separable
An audit of 254 SWE-bench submissions across four splits, done without running any models, finds the top ten agents share 285 successes and 51 failures, leaving only 164 instances that distinguish them at all. Exact paired McNemar tests separate none of the 29 adjacent Verified top-thirty pairs at alpha=0.05, while within-model scaffold ranges reach 29.8 percentage points against an 8.8-point spread across the entire top thirty. For a builder picking a coding agent, the scaffold you wrap it in matters roughly three times more than the leaderboard rank, and the paper releases a five-step audit protocol plus the instance partition so you can run the comparison on your own task set.
↳ Follow the thread