'How Do Agents Fail on AutoResearch': 800 Trajectories Across 8 Harness-Model Combos Find One Failure Every Model Shares — They Never Check Output Against Evidence
AutoResearchEval (arXiv 2608.14905, submitted 2026-08-14) covers 100 real frontier research tasks across seven scientific domains and the full lifecycle — ideation, retrieval, execution, analysis, writing, review — and annotates 800 agent trajectories into a 45-pattern AutoResearch Failure Taxonomy. The headline result is not a leaderboard but a shared deficit: agents lack "the ability to check what they produced against what they found, revise when it does not hold up, and question whether the path they took was sound," and this held across all 8 harness-model combinations including the strongest models tested. That points at a model-level limitation, not a scaffolding one — building a better harness will not patch it.
↳ Follow the thread