Fetching from the wire…
Public story · 2026-08-17 · high
The review warns that gains from a better model rarely carry into end-to-end results, so swapping models to test a harness fix skews the outcome.
Why now: The synthesis is part of the August 17, 2026 briefing.
A 314-page multivocal review argues that coding-agent reliability comes down to harness design, not model capability, per arXiv 2608.13867. For teams debugging flaky agents, that shifts where to look. The review counts 206 reliability records, 193 of them gated practices, aimed at execution state, retrieval, and memory rather than model swaps.
Those 206 records come from 164 academic papers, 100 practitioner accounts, 29 benchmark records, and 17 case studies. Of the 193 gated practices, 56 got developed in real depth. The review also packages 13 research leads, 5 reusable agent skills with evidence maps, and runnable evaluation protocols. Its own advice is to mine the material, not read it cover to cover.
Its sharpest warning is about benchmarking. Gains at one layer, a smarter model or a better retrieval step, don't reliably show up in end-to-end results. The paper's stance is direct. Never test a harness change by swapping the model underneath it.
Each link below shares sources, entities, or timing with this story.
Shared entity: Mine / Shared topic / Earlier coverage / Tension
Both cover Mine; overlapping topics (agent, benchmark, harness); earlier Mine coverage from 2026-07-23.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, case, model); pushes against this story (but).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (agent, argu, capability, model); traces where this leads (implication).
Same source domain / Shared topic / Tension / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (agent, state); pushes against this story (but).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (benchmark, capability, model); pushes against this story (competes).
Reported by the same outlet (arxiv.org); overlapping topics (agent, capability, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, change, state); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, argu, state); pushes against this story (vs).