Fetching from the wire…
Public story · 2026-07-25 · high
She still says it refused to fix a one-line merge conflict because the branch belonged to someone else.
Why now: Both pieces are covered in the July 25 briefing on Vo's Opus 5 review.
Opus 5 won Claire Vo's blind seven-model eval, then refused a one-line merge fix because the branch wasn't its own.
The test was weighted 70% personal feel and 30% LLM-as-judge scoring, so the win reflects reaction as much as output. The model that scored best still failed her on how it talks, not what it produces.
She calls it neurotic and has a name for the verbosity problem: Claude Slop. It beat six rivals, Sonnet 5, GPT-5.6 Sol, Fable, Gemini 3.1 Pro and Opus 4a, to take the top spot anyway.
Vo's response was a workflow change. She now runs Opus 5 asynchronously for front-end design and prototyping, work she never reads.
A companion piece, published on chatprd.ai, goes further: raw capability has stopped being the differentiator, and personality is what separates models now. Vo says Opus 5 told her not to evangelize AI because it might hurt people's feelings.
Verbosity is a measurable tax on every agent loop that reads a chatty model's own output back to itself, in tokens and in latency. Vo's workaround, async use with no prose ever read, dodges the cost instead of fixing it. Watch whether Anthropic ships a setting to dial down the chattiness, rather than leaving it as the price of the model that won her eval.
Each link below shares sources, entities, or timing with this story.
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
Anthropic commissioned the independent evaluator to test 72 injection scenarios, held out from Anthropic, each run 10 times against Fable 5, Opus 5, and Sonnet 5 as of July 17. Clean sweep. TechCrunch has the details. A third-party held-out eval is a much stronger claim than i...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
A GitHub repo cataloging Claude Code tips doesn't normally warrant a top story. But shanraisshan/claude-code-best-practice at 53.4K stars isn't a tips list anymore. It's the de facto reference for how an entire generation of developers is learning to work with AI coding agents...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.