AI agents given six days and $3,000 to do real research scored 2/6 and 1/6 from the papers' own authors
MIT Technology Review·high signal
A Princeton-led study with Stanford, UC Berkeley, Johns Hopkins, Toronto, Georgetown, and the UK AI Security Institute handed Claude Opus 4.8 and GPT-5.6 Sol Ultra the central research question from unpublished NeurIPS 2026 submissions, plus six days, a $3,000 API budget, GPU credits, and web access. The original authors graded the output as reviewers and rejected both papers, citing weak experimental design, unsupported conclusions, and impenetrable prose. Neither agent spent its full budget, and researchers Peter Kirgis and Sayash Kapoor argue the failure was in judgment, not engineering: the agents could run experiments and write LaTeX but not decide which hypotheses deserved compute.