Fetching from the wire…
Public story · 2026-08-17 · high
The study scored seven frontier models on 36 R&D tasks using three new process metrics, not pass/fail outcomes alone.
Why now: The paper's break from outcome-only benchmark scoring is why it's in the August 17 briefing.
Seven frontier models tackled 36 long-horizon R&D tasks, scored on how they worked rather than whether they finished, per arXiv 2608.13417.
That matters for anyone building or evaluating agent harnesses: 36 tasks across seven models is enough to expose a pattern, not a fluke. The paper argues the inconsistency traces to scaffolding choices a team controls, not to the model itself.
A 13-author paper drops outcome-only scoring for three rule-based metrics: Solution Framing, Execution and Feedback Control. Every model framed a workable approach and executed it, but none repeated the same approach twice.
Results swung run to run. Models adapted known techniques instead of inventing new ones, and every model stalled at identifiable points in the process rather than failing at random. The paper traces that swing to three sources: experience reuse, harness design and model stability.
None of those three sources live inside model weights. They live in the scaffolding around the model. A separate study on agent skills reaches the same conclusion from a different angle, the paper notes.
The real test comes next: rerun these seven models with better experience reuse and steadier harnesses, and see whether the run-to-run swings shrink. If they don't, the paper's own metrics point somewhere else entirely.
Each link below shares sources, entities, or timing with this story.
Shared entity: Execution / Same source domain / Shared topic / Earlier coverage
Both cover Execution; reported by the same outlet (arxiv.org); overlapping topics (agent, author).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (agent, bottleneck, control, harness, model).
Reported by the same outlet (arxiv.org); overlapping topics (agent, call, harness, model, paper).
Shared entity: Execution / Shared topic / Earlier coverage
Both cover Execution; overlapping topics (agent, harness, model); earlier Execution coverage from 2026-05-14.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, call); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, author, benchmark); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (agent, call, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, author, call); pushes against this story (against).