Research
CryptanalysisBench: Frontier Models Break 65–86% of Known-Broken Ciphers and Find Novel Attacks
A 191-task benchmark across six families of cryptographic primitives, evaluated on Claude Opus 4.8, Sonnet 5, Mythos 5, GPT 5.5, and GLM 5.2. Models broke 65–86% of schemes with known attacks, cracked 6–12 previously unbroken schemes at full strength, and handled 24–61 scaled-down variants. Most significantly, the models produced genuinely novel results — a key-recovery vulnerability in SpoC AEAD and an error in KINDI's security proof — putting LLM cryptanalysis at the edge of published state of the art rather than merely reproducing textbook attacks.
↳ Follow the thread