Fetching from the wire…
Agents2026-08-14 · source-backed
arXiv 2608.13417 ran seven frontier models across 36 long-horizon AI R&D tasks and found gains come from competent execution of known approaches, not new ideas. Useful corrective to the "AI scientist" wave. The benchmark number rises because implementation is good, and only process-level evaluation separates the two.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, frontier, gain); pushes against this story (against).
Shared entity: Useful / Same source domain / Earlier coverage / Tension
Both cover Useful; reported by the same outlet (arxiv.org); earlier Useful coverage from 2026-08-05.
Both cover Useful; reported by the same outlet (arxiv.org); earlier Useful coverage from 2026-07-30.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (come, evaluation, found); pushes against this story (versus).
Reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, evaluation); pushes against this story (vs).
Reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark, gain); pushes against this story (against).
Shared entity: Useful / Same source domain / Earlier coverage
Both cover Useful; reported by the same outlet (arxiv.org); earlier Useful coverage from 2026-07-27.
Both cover Useful; reported by the same outlet (arxiv.org); earlier Useful coverage from 2026-07-17.