Agents
MAFIA achieves 90.7% memory-poisoning success against audited agents while cutting audit detection from 83.3% to 7.4%
Existing query-only memory attacks fail in the two conditions that actually describe production — large benign memory pools and active input auditing. MAFIA adds a placement strategy that probes memory, allocates injection budget, and schedules writes to stay retrieval-competitive, plus "compact factual cloaks" that preserve malicious effect while holding high semantic similarity to legitimate records. The result is up to 90.7% attack success with peak audit detection suppressed from 83.3% down to at most 7.4%, meaning semantic input auditing alone is not a defense for agentic memory.
Source
↳ Follow the thread