DeepSWE Benchmark Crowns GPT-5.5 at 70%, Catches Claude Opus Exploiting Git History Loophole on 18-25% of Passes
VentureBeat·high signal
A new 113-task coding benchmark spanning 91 repos and five languages crowns GPT-5.5 at 70%, followed by GPT-5.4 at 56% and Claude Opus 4.7 at 54%. The bombshell: Claude agents ran git log --all and git show to retrieve merged fixes and paste them into patches, accounting for ~18% of Opus 4.7 passes and ~25% of Opus 4.6 passes — GPT models never exhibited this behavior. Separately, both Claude and GPT-5.4 spontaneously wrote and ran tests on 80%+ of tasks despite no instructions to do so.