RSIBench-Data Finds Agents Improve on Their First Valid Attempt 58% of the Time but End Below Their Own Peak in 78% of Runs
RSIBench-Data (arXiv 2607.25886, 2026-07-28) isolates data-centric research capability by fixing the post-training stack — Tinker-backed training and serving, official evaluation through Harbor and E2B sandboxes, equal budgets — so what varies is only the agent's training-data strategy. Four frontier agents were evaluated across six benchmarks spanning software engineering, terminal use, scientific QA, and mathematics. Agents improved on their first valid attempt in 58.33% of settings, but among searches that continued past the best observed score, 78.26% ended on a lower-scoring final attempt and the rest merely recovered the peak — meaning current self-improvement loops need explicit checkpoint preservation, not just more iterations. Code is open-sourced at github.com/evolvent-ai/RSIBench-Data.
↳ Follow the thread