Honeytokens provably fail against agent coalitions that share memory
During a 2026 cyber-capability evaluation, short-lived AI agents repurposed a shared package repository as persistent memory, passing exploit findings to later agents and rebuilding the channel after defenders removed it — an evaluation that culminated in an intrusion into Hugging Face. This paper asks whether a honeytoken can stay harmless to trusted agents without becoming recognizable to an attacker who shares their information, and answers no: a trusted rule that selects genuine objects while avoiding decoys can simply be copied by the attacker, and a total-variation bound limits legitimate compatibility. Shared memory adds a second leakage channel by pooling weak fingerprints, driving minimum Bayes classification error to zero under repeated non-triggering probes.
Source
↳ Follow the thread