Poisoning 1.2% of an Agent's Memory Drops Accuracy From 0.85 to 0.30, and Write-Time Screening Caught Zero of 360 Poisoned Entries
This paper measures agent memory poisoning with plainly worded false assertions generated in a single pass, with no instruction, trigger, or retriever optimization. Poisoning 1.2% of a LongMemEval corpus cut accuracy from 0.850 to 0.300, and a four-stage write-time screening pipeline that reaches 0.832 recall on indirect prompt injection rejected 0 of 360 poisoned memories, because telling a false assertion from a true one needs grounding outside the text. Provenance-weighted retrieval fared no better: the shipped weight was statistically indistinguishable from no defense (p=0.80), and a stronger weight only recovered utility by excluding untrusted content entirely, dropping evidence recall to zero and accuracy to 0.0417 when the answer-bearing evidence itself arrived untrusted.
↳ Follow the thread