Agents
Seven frontier models on 36 long-horizon research tasks: agents implement competently, but methodological novelty stays rare
arXiv 2608.13417 (2026-08-13) evaluates seven frontier models across 36 long-horizon AI R&D tasks and argues that final scores hide what agents are actually doing. Agents reliably formulate and implement practical solutions, but genuine methodological novelty remains rare — the gains come from competent execution of known approaches, not from new ideas. It is a useful corrective to the wave of 'AI scientist' claims: the benchmark number goes up because the implementation is good, and process-level evaluation is what separates the two.
Source
↳ Follow the thread