Re-run an effort sweep on Opus 5 — and stop expecting lower effort to shorten responses
Anthropic's guidance for Opus 5 reverses the Opus 4.7/4.8 advice of "start at `xhigh` for coding": start at the default `high`, "use `low` and `medium` liberally as your primary control for token cost and response time wherever your evals show quality holds," and explicitly "run a fresh effort sweep on your evals rather than reusing them" if you carried settings over. The non-obvious trap is that effort controls thinking volume, not visible output — "changing effort does not reliably shorten responses," so teams lowering effort to cut verbosity are paying a capability tax for nothing and should prompt for length instead. Effort also shapes tool-call behavior: lower levels combine operations into fewer calls and skip preamble, higher levels make more calls and explain plans first.
↳ Follow the thread