Fetching from the wire…
Public story · 2026-09-26 · high
Terminal-Bench runs show token use triples from low to max effort, buying more verification, not a better plan.
Why now: Shihipar published the workflow on September 25.
Claude Code's effort setting mostly buys extra verification, not a smarter plan, per Shihipar's workflow for spending effort. That matters for token budgets. On Terminal-Bench 3.0, the same task took Fable 5.1 a median 73k tokens at low effort and 222k at max, a 3x jump.
The extra tokens buy real accuracy. Security task accuracy was 64% at low effort and 87% at high. Hardware task accuracy was 34% at low and 75% at high.
The workflow starts with an interview against a spec. Shihipar has Claude implement and iterate on low effort, then switches to high effort for verification and tests once the code exists.
If a task still fails at high effort, the fix is a better spec. Turning the dial to max just spends more tokens hunting for an approach that was never going to work.
The post doesn't say how this holds outside Terminal-Bench 3.0, or whether it applies to tasks that aren't security or hardware focused.
Each link below shares sources, entities, or timing with this story.
The reflex is to crank effort to max when a task is hard. Thariq Shihipar's post on claude.dev, published September 25, argues that reflex is wrong, and it costs you 3x in tokens for a benefit you probably didn't want. The finding is that effort controls verification and edge-...
The model built its own benchmarks and shipped over 3,000 merged changes across 150-plus concurrent threads with no customer-facing incidents.
At Rails World 2026, David Heinemeier Hansson said agent-generated code is now the default at 37signals, and hand-writing code is the exception kept for fixing the workflow when an agent misfires. He put a number on his own month: about 150,000 lines of production code in Augu...
3,000+ merged changes. 150+ concurrent threads. No customer-facing incidents. P75 web fresh load went from 3,085ms to 550ms. Desktop cold start, 6,310ms to 3,328ms. Sending a message in Cowork cloud, 928ms to 48ms. Average of 3.1x faster (Claude blog, published September 23)....
The technique filters terminal output an agent reads one line of, and resolution rates held steady on 50 SWE-bench Lite tasks.
A rival's harness beat the model's own vendor harness in 9 of 12 pairings, and Claude Code's starting context ran over 10x Pi's before work began.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.