Fetching from the wire…
Public story · 2026-07-30 · high
Claude-Fable-5 scored 56.4% on accounting rubrics graded by real accountants, but its work rarely held up end to end.
Why now: The benchmark surfaced in the July 30, 2026 briefing, and the pass-rate gap is the detail worth sitting with before anyone treats the rubric score as a readiness signal.
Mercor and Ramp built an accounting benchmark graded by real accountants, and no model cleared a 2.6% Pass@8 completion rate, per the paper on arXiv.
That gap is the number to sit with if you're weighing whether an agent can close your books unsupervised. The benchmark's best rubric score, 56.4%, sits nowhere near a rate you'd trust without a human checking the work.
The benchmark, called APEX-Accounting, spans 160 private tasks across 10 self-contained worlds, each with its own accounting system, spreadsheets, and PDFs. Practicing accountants wrote every task, solved it themselves, and built the grading rubric, so there's no ambiguity about what counts as done.
On that rubric, Claude-Fable-5 running at Max effort led with 56.4% Mean Criteria@3. Muse-Spark-1.1 at xHigh effort followed at 52.6%. Check Pass@8, the rate at which a model finishes a task well enough to actually use, and both scores collapse to under 2.6%.
The paper also documents a Simpson's paradox in token spending. Raise the budget from $1 to $50 per task and the aggregate score climbs, but individual tasks often score worse with more tokens spent. APEX-Accounting doesn't say why. My read: more tokens just give a model more room to talk itself out of a right answer.
56.4% on a rubric means the model understood the assignment. It doesn't mean the ledger balances. Watch whether the next accounting benchmark leads with Pass@8 instead of rubric credit. That's the signal that builders took this gap seriously.
Each link below shares sources, entities, or timing with this story.
The Pragmatic Engineer published a deep read on August 25 of Inspect, the coding agent Ramp built instead of standardizing on Claude Code or Cursor. The numbers: Inspect authors 75% of Ramp's merged PRs, 90% of PRs in its own repository, passed 1 million total sessions in July...
Anthropic commissioned the independent evaluator to test 72 injection scenarios, held out from Anthropic, each run 10 times against Fable 5, Opus 5, and Sonnet 5 as of July 17. Clean sweep. TechCrunch has the details. A third-party held-out eval is a much stronger claim than i...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
This one's been building for days and it crystallized this week. Per The Register, the incident behind the US export-control block on Anthropic's Fable 5 and Mythos 5 wasn't a jailbreak or a guardrail bypass. It was a plain three-word prompt, "fix this code," run against CVE-l...
An automated framework evaluated GPT, Gemini, Claude and Grok on 85 algorithmic C# tasks derived from HumanEval, producing 340 solutions scored on three independent axes: functional correctness via unit tests, static quality via Roslyn AST analysis, and runtime efficiency via...
A $30B software vendor is trading headcount for tokens. That's the story. 404 Media obtained an internal SAP email dated July 1 that suspends most travel and new hiring, citing "rising token usage and costs as more AI-driven scenarios go live" as the reason. Exceptions carved...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.