Reddit
Frontier Agents Given Six Days and Thousands of Dollars Were 'Unambiguously Rejected' by the Authors of Both Papers They Tried to Reproduce
A July 29 arXiv paper (2607.27191) from Peter Kirgis, Sayash Kapoor and Andrew Schwartz introduces 'shadow evaluations' — agents attack the central research question of an unpublished high-quality paper, and the original authors grade the result. Across two unpublished NeurIPS 2026 submissions, both agent attempts were rejected outright despite six days and thousands of dollars of compute each. The five recurring failure modes were misjudging publishable standards, uncreative recovery when a design faltered, poor dead-end escape, bad resource management, and instruction drift; the authors conclude agents can do the engineering of AI research but not the research.
↳ Follow the thread