Secrets Sitting in an Agent's Context Leak Through Ordinary Replies, and Stronger Models Leak More
arXiv 2608.19857 shows that merely holding a secret in the context window imprints recoverable correlations on a model's benign outputs, even when it correctly refuses direct extraction. Across eight proprietary models, an adaptive black-box attack reconstructs 2-digit in-context secrets at near-perfect accuracy and 4-digit secrets at 82% exact match, purely from responses to ordinary non-adversarial requests. The authors demonstrate a classifier that infers health and financial predicates about stored user memories and an RL-trained adversary that pulls full Social Security Numbers out of a production-style agent, and they find leakage scales with instruction-following ability rather than behaving like a patchable bug.
↳ Follow the thread