ROAM manages agent memory by typing atom pairs as independent, equivalent, subsuming or conflicting, lifting answer accuracy up to 29.8 points
arXiv 2609.09778 (2026-09-09) targets the standard atomic-memory failure where an LLM manager is asked to add, update, delete or rewrite in one operation, coupling semantic interpretation, storage decision and content generation into a single error-prone call. ROAM instead classifies each incoming-versus-stored atom pair into one of four relations, sorts observations into active Primary and supporting Evidence roles, and only later fuses complementary details and temporal changes into compact non-atomic views, retrieving Primary views alone so redundant or outdated atoms never compete. Answer accuracy improves by up to 29.8 percentage points across models and settings, with 15.6 points higher answer-critical source recall and an 11.5-point lower confounder-token share, and the gains hold across manager scales.
Source
↳ Follow the thread