CIPL measures what an outside observer can actually recover from an agent, not what its memory contains
arXiv 2609.21686 (18 Sep) argues that privacy leakage in LLM agents is usually evaluated inside one component (memory, retrieval, tool pipeline), which conflates internal exposure with information an attacker can reconstruct. CIPL models a target through sensitive source, selection, assembly, execution, observation and extraction stages and scores the transition from selected sensitive units to attacker-recoverable output under one protocol. Across memory-based, retrieval-mediated and tool-mediated targets plus a live BrowserUse case study, storage labels alone did not determine recoverability: memory targets are near-saturated, retrieval leakage is frequently partial, and tool-mediated and live-agent leakage swings with observation surface, prompt-to-channel alignment, retrieval depth and provider behavior, with a stratified semantic audit catching disclosures exact matching misses.
Source
↳ Follow the thread