Fetching from the wire…
Public story · 2026-07-27 · high
An independent GitHub writeup found the gain came from writing five times more code, then flagging 93% of it as slop.
Why now: The score comes from a July 27 read of the independent writeup, not from Anthropic or the benchmark's own authors.
Opus 5 passed 24% of a 17-checkpoint coding benchmark called SlopCodeBench, quadruple the 6% that Sonnet 5 and Opus 4.8 each managed, per an independent writeup posted to GitHub.
That's the number teams will see first if they're picking a model by benchmark score. What they won't see up top is that Opus 5 flagged 93% of its own new code as slop to get there.
Strict pass on this benchmark means clearing every new test a checkpoint adds, plus every regression test inherited from earlier checkpoints. Opus 5 cleared four of seventeen checkpoints; Sonnet 5 and Opus 4.8 each cleared one. The paper's published baseline for Opus 4.6 sat at 17%, so Opus 5 beat that too.
SlopCodeBench doesn't hand a model the full spec up front. It reveals requirements one checkpoint at a time instead.
Opus 5 wrote roughly five times more functions than Sonnet 5 and Opus 4.8. Test code made up 51% of its total output, against 11-24% for the other two models.
So the model didn't get leaner. It got more prolific, writing more tests to catch its own bugs.
Whether that trade is worth it depends on what you value in production code, and the writeup leaves the call to the reader.
Each link below shares sources, entities, or timing with this story.
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Opus); both cover Opus, Sonnet; overlapping topics (benchmark, code, opus).
Claude Code uses Opus / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Opus); both cover Opus, Sonnet; reported by the same outlet (github.com).
Claude Code uses Opus / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code uses Opus); both cover Opus, Sonnet; overlapping topics (benchmark, code, test).
Linked by a graph relationship (Claude Code uses Opus); both cover Opus, Sonnet; overlapping topics (code, test).
Linked by a graph relationship (Claude Code uses Opus); both cover Opus, Sonnet; overlapping topics (code, opus).
Claude Code uses Opus / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Claude Code uses Opus); both cover Opus, Sonnet; reported by the same outlet (github.com).
Opus built by Anthropic / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Opus built by Anthropic); both cover Opus, Sonnet; overlapping topics (code, opus).
Claude Code uses Opus / Shared entity: Opus / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code uses Opus); both cover Opus; overlapping topics (benchmark, code, opus).