Research
AIDE^2 Rewrites Its Own Research-Agent Code for 8 Days and Keeps Seven Improvements That Transfer to Four Held-Out Benchmarks
AIDE^2 proposes changes to its own code, benchmarks the modified agents on AI R&D tasks, and keeps only the changes that win on hidden evaluations. In an autonomous 8-day run it found seven successive improvements, including a new search policy and context-compressing memory mechanisms. The gains generalized to held-out ML engineering, heuristic algorithm engineering and weather-forecasting benchmarks. This is a concrete, evaluated instance of recursive self-improvement, gated on hidden evals rather than self-reported scores.
Source
↳ Follow the thread