Fetching from the wire…
Top 5 · 2026-06-16 · source-backed
Here's the number that should change how you read every coding-agent ranking: 99 of 100 entries on the SWE-bench Verified leaderboard are vendor-submitted. One carries an independent verification badge. The other ninety-nine are labs grading their own homework on their own scaffolds. Digital Applied laid it out on June 16, and an arXiv paper the same week made the deeper point: benchmark scores tell you what an agent got right, never how it got there.
The spread is the story. Claude Opus 4.5, the exact same weights, scored anywhere from 50.2% to 55.4% depending purely on how the agent system managed context and tool calls. Push across model versions and harnesses and you get 51.9% on Scale's setup versus 69.2% on Anthropic's. That's a 17 to 21 point swing with the model held roughly constant. The variable isn't intelligence. It's plumbing.
I've felt this directly. I rebuilt the scaffolding around a research agent in my pipeline. Same model, no prompt-engineering tricks, just better context windowing and a tighter tool-result format, and the pass rate moved more than any model upgrade had given me in months. At the time I assumed I'd gotten lucky on a few cases. Now I think that was the actual lesson and the model swaps were the placebo.
So why do we keep treating the leaderboard like a shopping list? Because "pick the model at the top" is a one-line decision and "profile your harness" is a week of unglamorous work. The leaderboard rewards the lazy read. One lab quietly stopped reporting SWE-bench entirely, which tells you they know the number isn't load-bearing.
What to do: stop A/B testing models before you've A/B tested your scaffold. Instrument your agent's trajectories, not just its final pass/fail. Where does it waste tool calls? When does it lose the thread in a long context? The arXiv "trajectories as programs" work argues you can fingerprint and even steer those patterns, which is a far better use of an afternoon than swapping opus for mythos in a config and hoping. And when a vendor quotes you a Pro score, ask which harness produced it. If the answer is "ours," it's marketing, not measurement.
Each link below shares sources, entities, or timing with this story.
Claude Opus built by Anthropic / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Opus built by Anthropic); both cover Claude Opus, SWE, Verified, When; overlapping topics (harness, leaderboard, model, opus, same).
Claude Opus built by Anthropic / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Claude Opus built by Anthropic); both cover Claude Opus, Same, SWE; reported by the same outlet (arxiv.org).
Boris Cherny works at Anthropic / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Boris Cherny works at Anthropic); both cover Anthropic, SWE, Verified, When; reported by the same outlet (arxiv.org).
Claude Opus built by Anthropic / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Claude Opus built by Anthropic); both cover Claude Opus, Same, SWE, Verified; overlapping topics (model, same).
Anthropic deprecates OpenClaw / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Anthropic deprecates OpenClaw); both cover Anthropic, SWE, When; reported by the same outlet (arxiv.org).
Anthropic released Fable / Shared entities / Same source / Shared topic / What happened next
Linked by a graph relationship (Anthropic released Fable); both cover Claude Opus, SWE, Verified; cite the same source (Digital Applied).
Claude Opus built by Anthropic / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Opus built by Anthropic); both cover Anthropic, Claude Opus, SWE, Verified; overlapping topics (model, number, swe-bench).
Anthropic partners with OpenAI / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Anthropic partners with OpenAI); both cover Anthropic, When; reported by the same outlet (arxiv.org, digitalapplied.com).