Vibe Coding
Pattern: Developers Are Rolling Back to Older Models Over Agent Behavior, Not Benchmark Scores
Opus 5 shipped July 24 near the top of SWE benchmarks (Snorkel AI ranked it second on Senior SWE-bench) yet drew a week of rollback posts — r/ClaudeAI's "Opus 5 is just annoying to work with. Back to Opus 4.8 for me" hit 338 upvotes, with Dan Shipper calling it "very hard to love" because it "argues with instructions, stopped before work was finished." Reported failure modes are scope expansion, overconfident certainty, and readier delegation that raises cost. Combined with the Gas Town postmortem, the signal is that instruction-adherence and stopping behavior now drive tool choice more than capability, and neither appears on a leaderboard.
Source
↳ Follow the thread