Fetching from the wire…
Agents2026-09-07 · source-backed
Long-running agents accrete compressed summaries, plaintext memory, pending tool plans and a KV cache, and today's "forget" deletes one plaintext record while every derived artifact survives (arXiv 2609.04875). Audited across three agent suites, nine baselines and three model families: memory deletion left leakage unchanged, instruction-based forgetting collapsed entirely under elicitation probes (Leak@probes = 1.00), and source redaction still acted on a revoked preference in 80% of episodes. Provenance-Guided Selective Replay, which crops the KV cache at the injection point and replays a sanitized suffix, matched a full reset at up to 9x fewer recomputed tokens. If you've told a user their data was deleted from an agent's memory, check which of those four artifacts you actually touched.
Each link below shares sources, entities, or timing with this story.
The attack needs no instruction, trigger, or retriever optimization, just plainly worded false assertions generated in one pass against a LongMemEval corpus. A four-stage screening pipeline that reaches 0.832 recall on indirect prompt injection rejected none of the poisoned me...
ROPE is structural rather than semantic: a value may reach a state-changing tool only if it traces unforgeably to the user, to a source the user explicitly named, or to the user's own authoritative records, checked deterministically over an audited set of sensitive parameters....
arXiv 2607.29167 describes the mechanism precisely: when an agent consolidates an external observation into long-term memory, the rewrite preserves the action trigger while erasing the low-trust source. The injected instruction resurfaces later looking like user history. Memor...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
Activation probes are usually evaluated against agents who don't know they're monitored, which is a generous assumption. This study held models, probes and thresholds fixed and varied only the disclosure: nothing, monitor present, or monitor present plus last round's score. Ac...
arXiv 2608.11632 argues storage retention doesn't identify *authoritative* state: unmediated updates by models, tools, and background workers cause stale overwrites and self-authorizing privilege escalation. Untrusted components propose typed changes against an exact predecess...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.