Fetching from the wire…
Public story · 2026-07-25 · high
Ethan Mollick found Opus 5 strong on short tasks but weaker on long ones, even as Musk calls it frontier-tier.
Why now: Techmeme grouped these four reactions together in the July 25 briefing, and no consensus has formed yet on how good Opus 5 actually is.
Reactions to Opus 5 split fast, with four prominent voices in AI landing on different reads of the same model, per Techmeme's roundup.
For builders choosing a model, the split points to two different things to optimize for: peak capability, or reliability across a long task.
Musk grouped Opus 5 with Grok 4.5, calling the pair alone on the Pareto frontier. Aaron Levie called it a huge jump over Opus 4.8.
Ethan Mollick pushed back with something more specific. He found Opus 5 quirky: strong on shorter tasks, less ambitious once the work stretches out.
Bindu Reddy split the difference, ranking Opus 5 behind Fable in sheer brilliance but still a strict improvement over 4.8.
Mollick's read is the one worth acting on. If Opus 5's weak spot really is long-horizon ambition, the fix is in how you hand it work, not in waiting for a better model.
Harnesses that break a big goal into short, verified steps should get more out of Opus 5. Ones that hand it one open-ended task probably won't.
Give Opus 5 a genuinely open-ended, multi-day task and see if the quirky label holds. Nobody's run that test yet.
Each link below shares sources, entities, or timing with this story.
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
Musk positioned it as trained on real Cursor sessions, landing fourth on the Artificial Analysis index, and Cursor's CEO called it a daily driver. But the top Hacker News thread centered on allegations Musk nudges outputs on political questions, while builders griped about "13...
Cursor's August 6 post details a two-stage router: "Compass" assigns each turn a 0–1 complexity score by predicting whether the user will be satisfied, then a taxonomy across domains (backend, frontend, database), tasks (bug fixes, commands, tests), and modifiers (visual chang...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
On July 9, Musk called Anthropic "obviously currently the leader in AI" and pledged never to cut off access, reversing his September 2025 claim that "winning was never in the set of possible outcomes for Anthropic." The context is structural: Anthropic pays roughly $1.25B/mont...
Two data points that tell the same story. First, Value Add Pulse counts four frontier launches in 30 days: Gemini 3.5 Pro, Grok 5, Anthropic's Fable 5 and Mythos 5, plus open-weight GLM-5.2 and Kimi K2.7. The model-layer moat compressed from quarters to weeks. Second, TechCrun...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.