Research
RECLAIM: The Best Agent Reproduces Only 41% of NeurIPS 2025 Papers Even When Code, Data and Weights Are Released
arXiv 2609.28850 fixes, per paper, the result to reproduce, the success criterion and a GPU-hour budget for 100 NeurIPS 2025 papers. An LLM grades runs from logs, not from the agents' own reports. The best of four agents reproduced 41% of Run-tier papers (code, data and weights released), 27% of Retrain-tier and 15% of Reimplement-tier. Failed runs used only 29% of their budget on average. The most common error, in 63 of 400 runs, was writing the method without checking any intermediate number against the paper.
Source
↳ Follow the thread