A web-browsing agent leaks recoverable signal about records it silently loaded, with no jailbreak and no disclosure
arXiv 2609.01663 formalizes context-inference attacks as a security game and evaluates them under decreasing attacker knowledge: a known context, an unknown context, and a context the agent retrieves through its own tool calls. A single unmodified attack carried through all three, reaching 100% ASR on small candidate sets and 63% at 1,024 candidates against a known context, 78.9 AUROC when the template and surrounding records are unknown, 92.5 AUROC when a 14B surrogate scores a 32B target, and 81.8 AUROC when records arrive as retrieval returns. The tested controls, an instruction not to disclose the context, logit suppression, and context dilution, did not stop it. Note the submission is dated 2026-08-31 though it surfaced in today's cs.CR listing.
↳ Follow the thread