Fetching from the wire…
Public story · 2026-03-18 · source-backed
A systematic comparison of 15 AI coding agents running the same Claude Opus 4.5 model found that Augment, Cursor, and Claude Code produced a 17-problem spread on 731 SWE-bench Verified issues. Not different models. Not different prompting strategies visible to the user. The same underlying model, producing meaningfully different results based entirely on how each product wraps it — context engineering, tool orchestration, and harness design. LogRocket
This empirically confirms what practitioners have suspected: the scaffolding around the model matters as much as the model itself. When you see a leaderboard score for "Claude Opus 4.5 on SWE-bench," you're actually seeing the score for a specific product's implementation of Claude Opus 4.5. Transfer that model to a different harness and you get a different number.
The implication for builders is concrete: benchmark against your specific codebase and toolchain before committing to a framework. A tool that scores highest on SWE-bench may not score highest on your repo's particular mix of languages, test patterns, and architectural conventions. The scaffolding gap means leaderboard results are a ceiling, not a guarantee. And a separate paper this week (see Research) found that 150 instances of the same model analyzing the same financial dataset produced substantially divergent conclusions — reinforcing that model-level capability is necessary but not sufficient for reliable outcomes.
Each link below shares sources, entities, or timing with this story.
Cursor supports Claude Opus / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Cursor supports Claude Opus); both cover Claude Opus, SWE, Verified, When; overlapping topics (harness, leaderboard, model, opus, same).
Cursor uses Opus / Shared entities / Shared topic / What happened next / Tension
Linked by a graph relationship (Cursor uses Opus); both cover Claude Code, Claude Opus, SWE, Verified; overlapping topics (claude, model, swe-bench).
Cursor benchmarked against Codex / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Cursor benchmarked against Codex); both cover Claude Opus, Cursor, SWE, Verified; overlapping topics (claude, model, opus, swe-bench).
Cursor benchmarked against Antigravity / Shared entities / Same source / Shared topic / Earlier coverage
Linked by a graph relationship (Cursor benchmarked against Antigravity); both cover Claude Code, Cursor, SWE; cite the same source (LogRocket).
Cursor benchmarked against Antigravity / Shared entities / Same source / Earlier coverage
Linked by a graph relationship (Cursor benchmarked against Antigravity); both cover Claude Code, Claude Opus, LogRocket, SWE; cite the same source (LogRocket).
Claude Code competes with Cursor / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Claude Code competes with Cursor); both cover Claude Code, SWE, Verified, When; overlapping topics (claude, model, same).
OpenCode competes with Cursor / Shared entities / Shared topic / What happened next
Linked by a graph relationship (OpenCode competes with Cursor); both cover Claude Code, Cursor, SWE, Verified; overlapping topics (claude, harness, model).
Cursor uses Opus / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Cursor uses Opus); both cover Claude Opus, SWE, Verified, When; overlapping topics (claude, opus, swe-bench).