UMPeek Recovers Private User Models From a Personalized Agent's Choices, With No Access to Memory or Backend
Personalized LLM agents increasingly compress retained memory into structured user models, and these are commonly treated as more privacy-preserving because direct memory-extraction attacks lose the source text they target. UMPeek shows the user model is itself a new attack surface: a black-box attacker forms hypotheses from choices a request leaves open, switches among ordinary follow-up tasks, and retains only claims supported and not contradicted by visible behavior, recovering private information purely from the personalized choices the model shapes. It outperforms existing attacks on both a benchmark across diverse personalization tasks and backends and in real-world systems using information confirmed to be retained, and the authors evaluate defenses against its adaptive probing.
↳ Follow the thread