Fetching from the wire…
Public story · 2026-08-03 · high
A new coding eval found that cranking reasoning effort higher made one model worse, not better.
Why now: Eval Suite 3 published August 1 and is covered in the August 3 briefing.
DeepSeek V4-Flash matched Grok 4.5's top coding score in Eval Suite 3, a benchmark of 23 frontier-hard tasks pulled from real merged open-source pull requests, for about a third of the price.
The eval, published by essamamdani.com on August 1, runs each task in a Docker sandbox and scores pass@1 with no retries and no majority voting, then logs the actual dollar cost per run. Three configurations tied at 91%: DeepSeek V4-Flash at medium effort ($2.04), Grok 4.5 at medium ($6.67), and Grok 4.5 at high ($13.16). If you're picking a model for a coding agent by price per correct patch, that gap is the whole decision.
The stranger result is what happened when effort went up. Grok 4.5 didn't improve from medium to high, it just cost twice as much for the same 91%. DeepSeek V4-Flash actually dropped, from 91% at medium effort to 87% at high. More reasoning tokens produced worse engineering decisions on these tasks, not better ones.
That's a testable claim, and a cheap one. If extra reasoning effort is quietly degrading your agent's PR quality while inflating your bill, that's worth an afternoon of checking against whatever eval or task set you already trust.
What the benchmark doesn't settle: it's one evaluator's run, coverage on Kimi K3 and Haiku 4.5 was partial, and there's no explanation offered for why more effort would hurt rather than help. Treat the top-line numbers as a snapshot, not a ranking to build a procurement decision on. The effort-tuning pattern is the part worth verifying yourself, since defaulting to higher reasoning effort is how most of us already configure these tools.
Each link below shares sources, entities, or timing with this story.
Kimi K3 competes with DeepSeek V4 / Shared entities / Same source / Shared topic
Linked by a graph relationship (Kimi K3 competes with DeepSeek V4); both cover DeepSeek V4, Flash, Grok, Test; cite the same source (Eval Suite 3).
Kimi K3 benchmarked against Fable / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Kimi K3 benchmarked against Fable); both cover DeepSeek V4, Flash; overlapping topics (cost, deepseek).
Kimi K3 competes with DeepSeek V4 / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Kimi K3 competes with DeepSeek V4); both cover DeepSeek V4, Flash; overlapping topics (cost, task).
Docker supports Claude Code / Shared entity: Grok / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Docker supports Claude Code); both cover Grok; overlapping topics (cost, decision, grok, task).
Kimi K3 competes with OpenAI / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Kimi K3 competes with OpenAI); both cover DeepSeek V4, Kimi K3; overlapping topics (cost, deepseek).
Docker supports Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Docker supports Claude Code); both cover Flash, Haiku; overlapping topics (deepseek, task).
Docker supports Claude Code / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Docker supports Claude Code); both cover Flash, PRs; earlier Flash coverage from 2026-07-23.
Docker supports Claude Code / Shared entity: Haiku / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Docker supports Claude Code); both cover Haiku; overlapping topics (between, cost).