Fetching from the wire…
Public story · 2026-07-30 · high
Anthropic built the benchmark with partners in Switzerland, Germany, and Israel; the models still found a real SpoC AEAD attack.
Why now: The tier breakdown is dated July 30, capturing where five frontier models stand on cryptanalysis right now.
Anthropic and four university partners built a 191-task cryptanalysis benchmark, and frontier models topped out at 86% only on ciphers that are already broken.
That gap matters for anyone taking AI capability claims at face value. On the 142 schemes with no public break, every model scored under 9% at full strength, per the benchmark site.
The test spans six primitive families, drawn mostly from AES, SHA-3, Lightweight Cryptography, and the Post-Quantum NIST competitions. Tier 1 holds 49 already-broken schemes. Tier 2 holds the 142 that are still standing.
Five frontier models ran against both tiers. Claude Mythos 5 topped Tier 1 at 85.7%, the strongest of the five, according to the benchmark site.
The models weren't empty-handed on the hard tier, though. They produced a genuine key-recovery attack on the full SpoC AEAD cipher and caught an error in the published CCA-security proof for KINDI, per the benchmark's authors.
The 86% headline score is recall of public write-ups, not cryptanalysis skill. The real number is the under-9% score against ciphers nobody's broken yet.
The tier breakdown is dated July 30, capturing where five frontier models stand on cryptanalysis right now.
Each link below shares sources, entities, or timing with this story.
CryptanalysisBench is 191 tasks across six families of cryptographic primitives, evaluated on Claude Opus 4.8, Sonnet 5, Mythos 5, GPT 5.5, and GLM 5.2 (arXiv 2607.18538). Models cracked 6–12 previously unbroken schemes at full strength and handled 24–61 scaled-down variants....
Anthropic's Frontier Red Team published a previously unknown attack on a NIST post-quantum signature candidate, found by a multi-agent system with a human collaborator who is not a lattice cryptography specialist. Claude also produced "Möbius Bridge," a fingerprinting techniqu...
Anthropic researchers used Claude Mythos to discover cryptographic weaknesses in the HAWK signature scheme and deliberately weakened AES variants. Human intervention was minimal and mostly consisted of encouraging the model not to give up and to push for publishable quality, n...
Green's July 29 assessment splits the two results sharply. He calls the HAWK attack significant precisely because HAWK was advancing toward standardization and "none of the ingredients are exotic," meaning the model won by thoroughly applying existing tools rather than inventi...
Timothy B. Lee reconstructs the sequence: page 13 of the Claude Fable 5 system card revealed Anthropic planned to silently degrade responses to prompts "targeting frontier LLM development," then switched after backlash to transparently downgrading such users to Opus 4.8. The r...
The most capable Claude tier is now government-gated. Let that sink in for a second. Not export-controlled to China, not restricted to enterprise. Gated by Commerce, for a U.S. company's U.S. customers. A June 26 letter from Commerce Secretary Howard Lutnick partially lifted t...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.