Fetching from the wire…
Public story · 2026-08-03 · high
A new coding eval found that cranking reasoning effort higher made one model worse, not better.
Why now: Eval Suite 3 published August 1 and is covered in the August 3 briefing.
DeepSeek V4-Flash matched Grok 4.5's top coding score in Eval Suite 3, a benchmark of 23 frontier-hard tasks pulled from real merged open-source pull requests, for about a third of the price.
The eval, published by essamamdani.com on August 1, runs each task in a Docker sandbox and scores pass@1 with no retries and no majority voting, then logs the actual dollar cost per run. Three configurations tied at 91%: DeepSeek V4-Flash at medium effort ($2.04), Grok 4.5 at medium ($6.67), and Grok 4.5 at high ($13.16). If you're picking a model for a coding agent by price per correct patch, that gap is the whole decision.
The stranger result is what happened when effort went up. Grok 4.5 didn't improve from medium to high, it just cost twice as much for the same 91%. DeepSeek V4-Flash actually dropped, from 91% at medium effort to 87% at high. More reasoning tokens produced worse engineering decisions on these tasks, not better ones.
That's a testable claim, and a cheap one. If extra reasoning effort is quietly degrading your agent's PR quality while inflating your bill, that's worth an afternoon of checking against whatever eval or task set you already trust.
What the benchmark doesn't settle: it's one evaluator's run, coverage on Kimi K3 and Haiku 4.5 was partial, and there's no explanation offered for why more effort would hurt rather than help. Treat the top-line numbers as a snapshot, not a ranking to build a procurement decision on. The effort-tuning pattern is the part worth verifying yourself, since defaulting to higher reasoning effort is how most of us already configure these tools.
Each link below shares sources, entities, or timing with this story.
VulcanBench Suite 3 found DeepSeek V4-Flash drops from 91% at medium effort to 87% at high, and Grok 4.5 is flat between medium and high while costing 2x. Run your own eval at two effort levels before defaulting to maximum. You may be paying double for worse decisions.
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
DeepSeek dropped V4 in mid-June as an open-weight model with a 1-million-token context window, priced at $1.74 per million input tokens, posting near-parity with GPT-5.4 on math and Q&A benchmarks (MindStudio). That's the headline number. The architecture underneath is more in...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
xAI launched Grok 4.5 and Grok Build on July 8, trained partly on Cursor developer-session data. The numbers are loud: 83.3% on Terminal-Bench 2.1, 64.7% on SWE-Bench Pro, priced at $2/$6 per million tokens. On a single coding task that works out to roughly $2.49 versus $11.80...
McKinsey's State of AI 2026 asked more than 1,700 respondents whether they skipped a software purchase because they could build it internally with agentic coding tools. 32% said yes. In technology it was 41%, among healthcare payers and providers 39% (CIO Dive). Temporal's Sta...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.