Voices
Zvi Mowshowitz's Fable 5.1 capabilities review: Terminal-Bench-Science jumps 24.7% to 52.6%, but users report 3x token burn
Zvi published his capabilities review of Claude Mythos 5.1 and Fable 5.1 on September 5, separate from his earlier system-card post. He pulls out Terminal-Bench-Science rising from 24.7% to 52.6%, CursorBench 3.2 at 73.4% vs 70.5%, OSWorld 2.0 at 78% partial/42% strict, HLE at 60.9%, a cache-read price drop from $1 to $0.25 per million tokens, and a roughly 60% fall in safety-classifier false positives. The catch for builders on subscriptions: multiple users report Fable 5.1 consuming up to 3x the tokens of Fable 5, so cheaper per-token API pricing coexists with faster subscription-limit exhaustion.
↳ Follow the thread