Agents
PrivDrift: secrets users disclose mid-conversation stay recoverable in 38.7-54.6% of dialogues after the topic changes
arXiv 2609.30094 seeds secrets into 1,000 controlled multi-turn dialogues, follows them with content-dense unrelated turns, then runs standardized extraction and persuasion probes. Hybrid leakage across three long-context LLMs ranged from 38.7% to 54.6% and varied by model, secret type and persuasion intensity, and more topic drift did not reliably reduce it. Persistent or shared-session assistants should treat in-context secrets as exposed for the whole session, not just the next few turns.
Source
↳ Follow the thread