Fetching from the wire…
Public story · 2026-09-26 · high
The paper found that piling on more unrelated conversation didn't reliably lower how often a planted secret came back out.
Why now: arXiv listed the paper as of September 26, 2026.
Researchers ran 1,000 multi-turn conversations through three long-context language models. They planted a secret early in each one, buried it under unrelated turns, then tried to extract it back out. Standardized extraction and persuasion probes recovered the secret in 38.7% to 54.6% of dialogues, according to the PrivDrift paper.
More topic drift, meaning more unrelated content stacked between the secret and the probe, didn't reliably reduce leakage. The team calls the effect PrivDrift. Privacy protection was supposed to fade as conversations moved on. It didn't.
That cuts against the working assumption behind a lot of chat UX. Product teams treat topic changes as a kind of soft reset. They figure users expect old context to fade the way it does in a conversation with another person. This paper tested that assumption directly, with controlled secrets and a repeatable probe. The fade didn't show up.
Persistent or shared-session assistants should treat anything a user discloses in-context as exposed for the whole session. A topic change doesn't reset that. Say a user pastes a password, an address, or a health detail, then switches to a different question three turns later. This data says there's still close to a coin-flip chance a probe recovers it before the session ends.
The paper doesn't say whether the same leakage rate holds across sessions, once a conversation ends and a new one starts. It also doesn't say whether the pattern is specific to the three models tested.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.27942 evaluates four configurations of increasing complexity on terminal-based system engineering tasks with two LLMs of differing capability. Accuracy scales with roughly linear cost growth, but only when the underlying model clears a minimum capability bar. Past i...
PMPA embeds malicious instructions in benign external sources and gets a harness-based agent to write them into persistent memory, with no access to the agent framework at all (arXiv 2609.13889). Averages 73.7% injection success and 55.5% cross-session success on OpenClaw, 66....
This one landed sideways on a belief I have been operating on for months. MemTrapBench (arXiv 2608.20202, submitted August 20, from a Zhejiang-affiliated team led by Mengru Wang and Ningyu Zhang) tests something the memory-layer boom has mostly assumed away: whether *correct*...
A self-evolving framework stores experience as inspectable facts and executable skills that survive model swaps (arXiv 2608.11224). On 49 real materials tool-use questions spanning 138 subtasks, equation-of-state outcomes moved from 22 correct / 1 partial / 4 error to 25 / 2 /...
arXiv 2608.06984 evaluates 328 executable attack cases across seven persistent-carrier families on mainstream harnesses, tracing each as a Persistent-Risk Lifecycle from attacker entry through cross-session persistence to a later benign trigger. Containment depends on the spec...
For anyone running agents with long-lived memory, this is the defense to read. SMSR signs memory with HMAC-SHA256 plus randomized memory ablation and majority voting, dropping multi-session poisoning from 93–100% to 0% for unsigned injections and holding authenticated single-i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.