Research
Memory-Based Self-Improving Agents Are Mostly Riding a Hidden Curriculum: Shuffle the Task Order and the Gains Shrink
A re-evaluation of two memory-bank self-improving agent methods (arXiv 2608.18066, 2026-08-18) adds two axes prior work skipped: multiple runs to quantify variance, and randomly shuffled task order. Both expose fragility. Evaluation noise in complex multi-step environments is already high, and stacking a self-improvement loop amplifies it; reported improvements depend heavily on the default task ordering, which acts as an implicit curriculum and a hidden prerequisite for success. Adding rubrics and environment feedback to memory construction, targeting the authors' underspecification hypothesis, only partially closes the degradation.
↳ Follow the thread