Skills
Frontier models now learn arbitrary ciphers from prompting alone, and encrypted harmful content slips past commercial classifiers as gibberish
Cipher-based covert-communication jailbreaks previously required fine-tuning a model on a corpus of encrypted harmful questions and answers. This work shows newer frontier models acquire the cipher through prompting and in-context learning, with alignment significantly weakened or entirely bypassed once communication runs through the learned encoding. Successful jailbreaks are demonstrated against models from Anthropic, Google and OpenAI, and the attack evades commercial harmfulness classifiers because the payload reads as nonsense text. Anyone running a content filter on agent inputs or outputs should assume string-level classification is not a control here.
↳ Follow the thread