Fetching from the wire…
Public story · 2026-08-26 · high
The split beat baseline memory on 35 of 37 model-benchmark pairs and cut common failure modes by as much as 80%.
Why now: The paper's own arXiv id, 2608.24876, dates it to August 2026.
Recuris, a memory architecture described in a paper posted to arXiv (2608.24876), splits an AI agent's memory into two separate stores. One tracks progress on the current task. The other holds skills the agent has learned over time. Most agent memory systems mix the two, so skill selection gets diluted by everything else in the agent's history.
On Opus 5, the split adds 15.6 points on tau-bench, a benchmark built around multi-step agent tasks. It adds 17.8 points on GPT-5.6 Sol, and on the longest-horizon tasks tested it reaches 32.2 points. Across 37 combinations of models and benchmarks, the approach won on 35.
The paper also adds a meta-agent that watches for failures and traces each one to a specific memory component, Working or Experiential. With the blame pinned to one store, common failure modes drop as much as 80%, instead of forcing a retrain of the whole system.
None of this needs a new model. It's a change to how memory gets organized, and it's cheap enough for any team building agent memory to test.
Each link below shares sources, entities, or timing with this story.
Claude Code benchmarked against GPT / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT, Opus; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT, Opus; reported by the same outlet (arxiv.org).
Claude Code benchmarked against GPT / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT; reported by the same outlet (arxiv.org).
Claude Code benchmarked against GPT / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT, Opus; overlapping topics (against, agent).
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT, Opus; overlapping topics (agent, gpt 5).
Claude Code benchmarked against GPT / Shared entity: Opus / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover Opus; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover Opus; reported by the same outlet (arxiv.org).