Fetching from the wire…
Public story · 2026-07-30 · high
The benchmark ran 310 test cases and found memory-backend choice swings repair success by 41.3 points but attack success by just 16.1.
Why now: As of July 30, MemSecBench is the clearest data point yet on how agent memory handles a poisoned entry after it's already in, not just whether it gets in.
Poisoned memory in AI agents persists 84.2% of the time, per MemSecBench.
That's not inert data sitting in a database. The full write-to-execute chain, poisoned memory turning into an actual action the agent takes, completes 50.3% of the time, per the paper.
MemSecBench ran 310 test cases across 48 contexts, per the arXiv paper. The matrix covered 24 configurations: two agent harnesses, four memory backends, three LLM backends.
Selective repair, stripping just the bad memory without breaking the rest, worked only 56.1% of the time. That number moved by 41.3 points depending on which configuration ran it, more than double the 16.1-point spread in attack success.
Vendors will keep selling memory backends on how hard they are to compromise. MemSecBench's numbers say the bigger difference between backends is how well they let you clean up after a breach, not whether one happens. Watch whether repair rate, not attack-resistance, becomes the number teams start asking vendors for.
Each link below shares sources, entities, or timing with this story.
LLM uses OpenAI / Same source domain / Shared topic / Tension
Linked by a graph relationship (LLM uses OpenAI); reported by the same outlet (arxiv.org); overlapping topics (agent, attack).
Simon Willison released LLM / Shared entity: LLM / Shared topic / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; overlapping topics (agent, chain).
LLM uses OpenAI / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; earlier LLM coverage from 2026-07-27.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-06-18.
LLM uses OpenAI / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (LLM uses OpenAI); both cover LLM; earlier LLM coverage from 2026-06-19.
Simon Willison released LLM / Shared entity: LLM / Earlier coverage
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-19.
Linked by a graph relationship (Simon Willison released LLM); both cover LLM; earlier LLM coverage from 2026-07-14.