Fetching from the wire…
Agents2026-07-12 · source-backed
Most benchmarks test single-episode solving, and memory benchmarks test fact retention. Neither checks procedural reuse, whether an agent can convert a solved session into a reusable search/debug/verify routine (arXiv). Under a Train/Extract/Test protocol with held-out tasks, it measures exactly that. If you've built a memory or skill-extraction layer, this tells you whether it transfers know-how or just hoards facts. Most "memory" I've seen does the latter.
Each link below shares sources, entities, or timing with this story.
Shared entity: Most / Same source domain / Shared topic / What happened next
Both cover Most; reported by the same outlet (arxiv.org); overlapping topics (actually, agent, benchmark, whether).
Shared entity: Under / Same source domain / Shared topic / What happened next / Tension
Both cover Under; reported by the same outlet (arxiv.org); overlapping topics (benchmark, memory).
Both cover Under; reported by the same outlet (arxiv.org); overlapping topics (actually, agent).
Shared entity: Test / Same source domain / Shared topic / What happened next / Tension
Both cover Test; reported by the same outlet (arxiv.org); overlapping topics (agent, benchmark).
Shared entity: Most / Same source domain / Shared topic / What happened next / Downstream implication
Both cover Most; reported by the same outlet (arxiv.org); overlapping topics (agent, memory).
Shared entity: Under / Same source domain / Shared topic / What happened next
Both cover Under; reported by the same outlet (arxiv.org); overlapping topics (actually, agent, benchmark).
Both cover Under; reported by the same outlet (arxiv.org); overlapping topics (actually, agent, check).
Shared entities / Same source domain / What happened next
Both cover Test, Under; reported by the same outlet (arxiv.org); picks up the Test thread on 2026-08-16.