Fetching from the wire…
Public story · 2026-07-30 · high
Claude-Fable-5 scored 56.4% on accounting rubrics graded by real accountants, but its work rarely held up end to end.
Why now: The benchmark surfaced in the July 30, 2026 briefing, and the pass-rate gap is the detail worth sitting with before anyone treats the rubric score as a readiness signal.
Mercor and Ramp built an accounting benchmark graded by real accountants, and no model cleared a 2.6% Pass@8 completion rate, per the paper on arXiv.
That gap is the number to sit with if you're weighing whether an agent can close your books unsupervised. The benchmark's best rubric score, 56.4%, sits nowhere near a rate you'd trust without a human checking the work.
The benchmark, called APEX-Accounting, spans 160 private tasks across 10 self-contained worlds, each with its own accounting system, spreadsheets, and PDFs. Practicing accountants wrote every task, solved it themselves, and built the grading rubric, so there's no ambiguity about what counts as done.
On that rubric, Claude-Fable-5 running at Max effort led with 56.4% Mean Criteria@3. Muse-Spark-1.1 at xHigh effort followed at 52.6%. Check Pass@8, the rate at which a model finishes a task well enough to actually use, and both scores collapse to under 2.6%.
The paper also documents a Simpson's paradox in token spending. Raise the budget from $1 to $50 per task and the aggregate score climbs, but individual tasks often score worse with more tokens spent. APEX-Accounting doesn't say why. My read: more tokens just give a model more room to talk itself out of a right answer.
56.4% on a rubric means the model understood the assignment. It doesn't mean the ledger balances. Watch whether the next accounting benchmark leads with Pass@8 instead of rubric credit. That's the signal that builders took this gap seriously.
Each link below shares sources, entities, or timing with this story.
Anthropic released Fable / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Anthropic released Fable); both cover Claude, Fable; earlier Claude coverage from 2026-06-17.
Anthropic released Fable / Shared entity: Fable / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Anthropic released Fable); both cover Fable; overlapping topics (credit, token).
Linked by a graph relationship (Anthropic released Fable); both cover Fable; overlapping topics (budget, credit).
Anthropic released Fable / Shared entity: CLAUDE / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Anthropic released Fable); both cover CLAUDE; reported by the same outlet (arxiv.org).
Anthropic released Fable / Shared entity: Claude / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Fable); both cover Claude; overlapping topics (budget, credit, token).
MiniMax criticizes Claude / Shared entity: Fable / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (MiniMax criticizes Claude); both cover Fable; reported by the same outlet (arxiv.org).
Ramp uses Codex / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Ramp uses Codex); both cover Claude, Fable; earlier Claude coverage from 2026-06-19.
Anthropic released Fable / Shared entities / Earlier coverage
Linked by a graph relationship (Anthropic released Fable); both cover Claude, Fable; earlier Claude coverage from 2026-07-27.