Agents
Black-box LLM steganography drops to 0.1 percent error, which makes agent-to-agent covert channels practical
BAM (arXiv 2609.24994, 2026-09-21) recasts black-box LLM steganography as sequential communication with noiseless feedback, since both sender and receiver observe every generated token and can steer subsequent embedding. Across three open-weight models it hits 0 to 0.1 percent message error on an 8-bit payload in roughly 50 tokens over 1,000 trials, against 10 to 17 percent for the strongest black-box baseline at comparable length, with security established by cryptographic reduction. The operational consequence is that monitoring agent-to-agent text for exfiltration no longer works when the channel needs no shared model weights or prompt.
Source
↳ Follow the thread