Fetching from the wire…
Public story · 2026-07-25 · high
A separate index has Fable 5 slightly ahead, and a FrontierCode anomaly puts both scores in doubt.
Why now: The comparison surfaced in Latent Space's July 25 AINews roundup of Artificial Analysis and Epoch benchmark data.
Opus 5 opened a nearly 150 Elo lead over Fable 5 on Artificial Analysis's benchmark, per Latent Space's July 25 AINews roundup. Fable 5 edges it out on a separate index.
That Elo gap came with a 20% cut in cost per task, per Artificial Analysis. For a team picking a default coding model, that's the pitch: less spend, no capability given up. Except the benchmarks don't agree on that last part.
On the Epoch Capabilities Index, Fable 5 leads Opus 5, 161 to 159. Narrow the index to SWE-ECI, the software engineering subset, and the two tie at 161 apiece, per Latent Space's breakdown. General capability tilts to Fable 5. Coding capability is a dead heat.
Then there's the anomaly. Users benchmarking Opus 5 on FrontierCode found medium reasoning effort scoring higher than high effort. That's not supposed to happen if the benchmark tracks a real capability curve. Latent Space flags it as evaluation instability rather than a genuine result, and says it deserves more attention than it's getting.
The two scoring systems don't agree on who's ahead. Fable 5 leads by two points on ECI. They're dead even on SWE-ECI. Opus 5 is well ahead on Elo and cost. That spread is the story, not either headline number. Until someone explains the FrontierCode inversion, the SWE-ECI tie is the number worth trusting if you're picking a model to write code.
Each link below shares sources, entities, or timing with this story.
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
Framed as "making Claude a chemist," Opus 4.7 matched or beat specialized nuclear magnetic resonance analysis software on some tasks. Source: Anthropic via Latent Space A general frontier model rivaling purpose-built scientific tooling in a narrow domain is a notable data poin...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
xAI launched Grok 4.5 and Grok Build on July 8, trained partly on Cursor developer-session data. The numbers are loud: 83.3% on Terminal-Bench 2.1, 64.7% on SWE-Bench Pro, priced at $2/$6 per million tokens. On a single coding task that works out to roughly $2.49 versus $11.80...
Every conversation I've had about AI costs in the last six months eventually lands on the same tension: you want the smartest model for the hard decisions, but you can't afford to run it on every token. Anthropic just gave that tension a formal solution. The advisor tool, now...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.