Checking only the expected output channel misses 46.9% of an agent session's real privacy exposure, and attacker self-reports are worse
ASLEval defines privacy exposure displacement, the gap between whatever local proxy you evaluate and the actual target-grounded exposure across a multi-step session. Its method pre-registers a hidden target set, then measures every declared visible exit rather than one designated action or final response, reserving internal traces for diagnosis. Across multiple enterprise-style environments and independently implemented runtimes, the expected-outlet-only view missed 46.9% of the exposure recovered by the union of visible exits, and attacker self-reports combined omissions with a high false discovery rate. Schema-aligned internal evidence usually preceded visible exposure at the request or probe level, giving an early-warning signal, and cutting model-visible returns changed the exfiltration path but could destroy normal-task success.
↳ Follow the thread