Agents
ScienceFlow hits a SOTA 70.22% Any-Medal on full MLE-bench inside a 24-hour budget by making research state rewindable
ScienceFlow (arXiv:2608.14354, submitted Aug 14, 2026, 19 authors) organizes long-horizon ML research into 'research segments' with recoverable executable states, and introduces ESTRA (Executable-State Transition through Re-Anchoring) to decide whether to continue a research direction or re-anchor on an active or archived state. It reports 70.22% Any-Medal on the full MLE-bench within a 24-hour compute budget, a 4.92-point improvement over the prior result. The builder takeaway: checkpointed, re-anchorable state is now beating pure retry loops on the hardest long-horizon agent benchmark.
Source
↳ Follow the thread