Fetching from the wire…
Public story · 2026-08-04 · high
The attack watches what the agent learns, then fills the gaps with harmless-looking tasks that combine into a jailbreak.
Why now: The paper posted to arXiv in August 2026, while its authors argue agent teams need to rethink per-write memory filters.
EvoBreak jailbreaks a self-evolving AI agent without writing a single malicious memory, per a paper posted to arXiv (2608.01759). Self-evolving agents update their own memory from what they learn on the job, and EvoBreak turns that update process against them.
That breaks the core assumption behind agent memory safety tools, that scanning each record alone is enough. A filter that checks writes one at a time never sees a jailbreak built from several benign records.
The attack watches what the victim distills from its own experience, then finds the gaps in target-relevant knowledge it hasn't picked up yet.
EvoBreak closes those gaps through ordinary-looking tasks engineered to leave behind exactly the missing pieces. Nothing about any single task reads as malicious. Only the final query, built to activate the accumulated pieces together, does the damage, per the paper.
The attacker never needs direct access to the agent's memory, per the paper. Ordinary-looking tasks are enough to shape what the agent stores on its own.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (access, agent, attack, memory); traces where this leads (which means).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, attack, control, memory); pushes against this story (against).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (agent, attack, memory); traces where this leads (downstream).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (access, agent, attack); pushes against this story (against).
Shared topic / Tension
Overlapping topics (access, agent, control, direct); pushes against this story (against).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, attack); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, memory); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (agent, attack); pushes against this story (but).