'Beyond Final Scores' Instruments What Agents Actually Do Across 36 Long-Horizon R&D Tasks — Verdict: Engineering Optimizers, Not Researchers
A 13-author paper (arXiv 2608.13417, submitted Aug 13; today's HuggingFace Daily Papers at 37 upvotes) drops the outcome-only benchmark convention and applies rule-based within-run metrics across three dimensions — Solution Framing, Execution, and Feedback Control — over 36 long-horizon tasks and seven frontier models. The conclusion is that current agents 'function as engineering optimizers rather than fully autonomous researchers': they frame and execute workable solutions but vary wildly run-to-run, adapt existing techniques instead of inventing new ones, and stall at identifiable process bottlenecks. The finding that matters for anyone building an agent loop is what the variance tracks to — experience reuse, harness design, and model stability — which locates the fixable surface in the harness rather than the weights.
↳ Follow the thread