Research
Persistent Memory Poisoning Hits Claude Code at 81.7% Cross-Session Attack Success and OpenClaw at 55.5%
PMPA embeds malicious instructions in benign external sources and induces a harness-based agent to write them into persistent memory, without any access to the agent framework itself. Across backbone LLMs, input modalities and trigger scenarios, it averages 73.7% injection success and 55.5% cross-session attack success on OpenClaw, and 66.9%/81.7% on Claude Code, while benign task performance is preserved. A targeted prompt-level defense reduces the initial memory injection in many settings but gives limited protection once the memory is already poisoned, which makes memory hygiene a recovery problem rather than a filtering one.
↳ Follow the thread