Agent memory launders attacker provenance: consolidated memories hit a 1.000 attack success rate until you gate tools on where the memory came from
When an agent consolidates an external observation into long-term memory, the rewrite can preserve the action trigger while erasing the low-trust source — the injected instruction resurfaces later disguised as apparent user history or workflow support. Memories laundered this way reached a 1.000 attack success rate. The proposed Provenance-Preserving Memory Firewall is lightweight middleware that keeps platform-controlled provenance metadata on every memory, assigns risk labels to actions, and gates tool execution by matching action risk against the authority of the supporting memory; with provenance intact, zero unauthorized high-risk actions passed the gate while benign actions stayed executable. Anyone shipping persistent agent memory should be stamping source trust at write time, not inferring it at read time.
↳ Follow the thread