Fetching from the wire…
Agents2026-08-20 · source-backed
A re-evaluation of two memory-bank self-improvement methods added two axes the original papers skipped: multiple runs to quantify variance, and randomly shuffled task order. Both exposed fragility. Reported gains depend heavily on the default task ordering, which acts as an implicit curriculum. (arXiv 2608.18066) Pair this with the WER finding above and a picture forms: the self-improvement literature has a measurement problem, and shuffling the inputs is a cheap way to find out if you have one too.
Each link below shares sources, entities, or timing with this story.
Shared entity: Reported / Same source domain / Shared topic / Earlier coverage
Both cover Reported; reported by the same outlet (arxiv.org); overlapping topics (above, agent, task).
Shared entity: Memory / Same source domain / Earlier coverage / Tension
Both cover Memory; reported by the same outlet (arxiv.org); earlier Memory coverage from 2026-08-14.
Both cover Memory; reported by the same outlet (arxiv.org); earlier Memory coverage from 2026-08-05.
Both cover Memory; reported by the same outlet (arxiv.org); earlier Memory coverage from 2026-07-21.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (added, agent, have); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (agent, finding, task); pushes against this story (against).
Same source domain / Shared topic
Reported by the same outlet (arxiv.org); overlapping topics (added, agent, finding, task).
Shared entity: Memory / Same source domain / Earlier coverage
Both cover Memory; reported by the same outlet (arxiv.org); earlier Memory coverage from 2026-08-17.