Fetching from the wire…
OSS2026-08-10 · source-backed
arXiv 2608.06984 evaluates 328 executable attack cases across seven persistent-carrier families on mainstream harnesses, tracing each as a Persistent-Risk Lifecycle from attacker entry through cross-session persistence to a later benign trigger. Containment depends on the specific carrier AND the harness-model pairing, and end-to-end attack-success rates hide where a chain actually breaks. If you run persistent memory or skill files, your safety posture is per-carrier. Hardening one does not generalize to the others. Pair this with the three-repos-no-format finding above and the picture is uncomfortable.
Each link below shares sources, entities, or timing with this story.
PMPA embeds malicious instructions in benign external sources and gets a harness-based agent to write them into persistent memory, with no access to the agent framework at all (arXiv 2609.13889). Averages 73.7% injection success and 55.5% cross-session success on OpenClaw, 66....
arXiv 2607.27942 evaluates four configurations of increasing complexity on terminal-based system engineering tasks with two LLMs of differing capability. Accuracy scales with roughly linear cost growth, but only when the underlying model clears a minimum capability bar. Past i...
For anyone running agents with long-lived memory, this is the defense to read. SMSR signs memory with HMAC-SHA256 plus randomized memory ablation and majority voting, dropping multi-session poisoning from 93–100% to 0% for unsigned injections and holding authenticated single-i...
The AEPD, Spain's data protection agency, disclosed an incident where a third party used an AI agent to autonomously chain a successful login, vulnerability discovery, access to personal data, and modification of invoices. No human stepped in between phases. The agency's frami...
This one landed sideways on a belief I have been operating on for months. MemTrapBench (arXiv 2608.20202, submitted August 20, from a Zhejiang-affiliated team led by Mengru Wang and Ningyu Zhang) tests something the memory-layer boom has mostly assumed away: whether *correct*...
Every skill marketplace runs on one assumption: certify each package, and the ecosystem is safe. CompoSkill breaks that assumption by showing composition risk is a path property, not a node property. The attack works black-box. The attacker knows only a role profile. They down...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.