Fetching from the wire…
Public story · 2026-09-06 · high
Fable 5.1's cache-read pricing fell 75% to $0.25 per million tokens, and safety-classifier false positives dropped about 60%.
Why now: Zvi posted the numbers on September 5, and the token-cost tradeoff he flags hasn't been settled since.
Fable 5.1's score on Terminal-Bench-Science more than doubled to 52.6%, according to Zvi's September 5 review of the model's benchmark results.
The jump matters for teams weighing whether the upgrade is worth a migration. Terminal-Bench-Science tests real terminal-task completion, not toy problems, and the model now clears it at more than double its earlier rate.
Other evals moved too. CursorBench 3.2 reached 73.4%, up from 70.5%. OSWorld 2.0 scored 78% on partial credit and 42% on the stricter pass, and HLE reached 60.9%.
Cache-read pricing dropped 75%, to $0.25 per million tokens, covering the reused context that agentic sessions rely on.
Safety-classifier false positives fell about 60%, meaning fewer benign actions get flagged mid-task.
Separate reports put Fable 5.1's token consumption at about 3x its predecessor per task, a cost that could offset the cheaper cache rate for high-volume workloads. Zvi's review doesn't reconcile that figure against the price cut for real-world usage.
Each link below shares sources, entities, or timing with this story.
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
The announcement describes the same underlying model at two safeguard levels: Fable generally available, Mythos restricted to vetted cybersecurity and life-sciences organizations, currently US-only. Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1 (against 24.7% for Fable...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
Anthropic published "Redeploying Fable 5" on July 18, and the headline reads like good news until you get to the metering. Fable 5 becomes a permanent subscription feature for Max and Team Premium. Good. But it's metered at 50% of standard weekly limits, meaning every Fable to...
Fable 5.1 came out this week and two people independently measured what it costs. They disagree by a factor of about four, and both are right. A MineBench run of 15 identical Minecraft builds put Fable 5.1 at $147.55 total against Fable 5's $54.93. Average inference time went...
Runta published FrontierHarness on September 2 and it's the most directly useful benchmark I've read this quarter, because it controls the one variable everyone conflates. Nine agent harnesses (Codex, Claude Code, OpenCode, Pi, Oh My Pi, DeepSeek Harness, Kimi Code, Exo Harnes...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.