Hacker News
Opus 5 Hits 24% on SlopCodeBench vs 6% for Opus 4.8 and Sonnet 5 — but Flags 93% of Its Own Lines as Slop
An independent benchmark writeup from the humanlayer context-engineering repo (113 points on HN) ran Opus 5, Opus 4.8 and Sonnet 5 against a 17-checkpoint subset of SlopCodeBench, a long-horizon benchmark released in March 2026 that reveals requirements progressively rather than upfront. Opus 5 scored 24% strict pass (4/17) against 6% (1/17) for both Opus 4.8 and Sonnet 5, and above the paper's 17% Opus 4.6 baseline — a strict pass requires all new tests plus every inherited regression test to pass. The caveat is the interesting part for builders: Opus 5 wrote 5× more functions than competitors, self-flagged 93% of its lines as "slop," and produced 51% test code versus 11–24% for the others.
↳ Follow the thread