MemGauge Finds Agent Memory Poisoning Risk Has a Threshold at Write Time and Couples to Utility at Read Time
MemGauge separately varies writing admission, management policy and retrieval exposure under matched clean and poisoned conditions, which prior work conflated by measuring utility or attack risk in isolation at fixed settings. Across 11 LLMs and two long-term memory benchmarks three distinct profiles emerge: a threshold-like risk transition during writing, policy-dependent local decoupling during management, and coupled growth of utility and risk during retrieval. Applying the same stage-level measurements to four existing memory systems produces consistent diagnostics, which means write admission is the stage where you can buy safety without paying utility, and retrieval is the stage where you cannot.
↳ Follow the thread