Agents
Coding-agent benchmark reality check: Codex CLI (GPT-5.5) tops Terminal-Bench 2.1 at 83.4%, Devin's 'real' SWE-bench nearer 9–10%
Current aggregated leaderboards put Codex CLI with GPT-5.5 at #1 on Terminal-Bench 2.1 (83.4%), Claude Code with Opus 4.8 at 78.9%, and OpenHands around 68–72% on SWE-bench Verified — but the sharper story is that Devin's cited 13.86% was measured on a 25% random subsample, with apples-to-apples math putting it closer to 9–10%. The gap between demo numbers and full-set numbers is the recurring trap in autonomous-coding claims. Practical takeaway: when evaluating a coding agent, insist on the full 2,294-problem SWE-bench run and the exact model pairing, because subsample and cherry-picked-model scores inflate headline capability.
Source
↳ Follow the thread