Fetching from the wire…
Public story · 2026-07-25 · high
Anthropic says lowering effort doesn't cut response length, only reasoning depth, and wants evals rerun from scratch.
Why now: Anthropic's effort docs now contradict the xhigh-for-coding advice teams adopted for Opus 4.7 and 4.8, leaving pinned evals unverified until rerun.
Anthropic reversed its coding guidance, telling developers to start Claude effort at high instead of xhigh, per its documentation. The stakes: effort controls how much Claude thinks, not how much it says. Per the docs, changing effort doesn't reliably shorten responses.
The new guidance frames low and medium as the main dials for cost and latency, not as fallback settings for when high runs too slow.
Anyone dropping effort to cut verbosity is paying for less reasoning and getting the same response length back, per the documentation. The fix Anthropic recommends is prompting for length directly instead of treating effort as a verbosity dial.
Effort also changes how Claude uses tools. At lower levels it combines operations into fewer calls and skips preamble before acting. At higher levels it makes more calls and explains its plan first.
Anthropic's own instruction is to run a fresh effort sweep on your evals rather than reuse old ones. Any production prompt still pinned to xhigh from the Opus 4.7 or 4.8 era should be treated as unverified until it's rerun against current behavior. The reversal undoes advice tied to those two releases, so anyone who set defaults then is running on guidance Anthropic no longer stands behind.
Each link below shares sources, entities, or timing with this story.
Official guidance separates two knobs people conflate constantly. Effort is not thinking time. It governs how many files Claude reads, how many tools it calls, and how many steps it takes before checking back. The rule: if Claude had all the context, clearly tried, and was sti...
Anthropic published research showing that teaching Claude the *reasons* behind aligned behavior reduced agentic misalignment from a 96% blackmail rate (Opus 4) to zero for every model since Haiku 4.5. A "difficult advice" dataset did it in 3M tokens vs. 30-85M for synthetic ap...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Platform release notes: Claude Opus 4.1 retires from the API on August 5. The experimental prompt tools APIs retire alongside the Workbench on August 17. The temporary 50% weekly usage boost for Claude Code subscribers runs through August 19. And Claude Sonnet 5's promotional...
Two weeks ago the US government forced Anthropic to pull Mythos 5 offline under an emergency export directive, on the theory that frontier cyber capability is dangerous enough to gate. This week a Chinese lab released a model you can download under an MIT license that benchmar...
The coding agent wars just entered a new phase. Cursor isn't just an IDE anymore. It's a model company. Cursor released Composer 2.5 on May 18 with a custom agentic coding model trained using 25x more synthetic tasks than Composer 2 and a novel "targeted textual feedback" appr...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.