Sources
Fable 5.1's headline gain is a science benchmark, not coding: 52.6% on Terminal-Bench-Science 0.1 against 29.0% for Opus 5
Simon Willison's 2026-09-01 writeup notes Anthropic's announcement spends most of its space on scientific research, citing 52.6% on the ten-day-old Terminal-Bench-Science 0.1 benchmark versus 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol, while the other benchmark movements are only slightly improved. Fable 5.1 exposes five reasoning levels (low, medium, high, xhigh, max) with no way to turn reasoning off entirely. Willison also flags that at effort low the transcript showed no summarized reasoning tokens at all, and that he had to fix llm-anthropic to record reasoning traces correctly before testing.
↳ Follow the thread