Fetching from the wire…
Public story · 2026-07-31 · high
It also found flaws in a deliberately weakened AES variant, with persistence, not capability, cited as the real limit.
Why now: Anthropic's account of the exercise became public on July 28, the clearest look yet at how long a model will stay on a problem with minimal supervision.
Claude Mythos spent about 60 hours finding cryptographic weaknesses in the HAWK signature scheme and a deliberately weakened AES variant, per Anthropic researchers.
For teams deciding how much autonomy to give AI systems on research work, this is the number that matters. The model didn't stall on the hard cryptography. It stalled on staying with the problem, and someone had to keep telling it to continue.
Human involvement was mostly encouragement, per the report. Researchers told the model not to give up and to push for publishable quality, not steering the technical work itself.
One caveat: the AES variant Claude found flaws in was deliberately weakened for the exercise, not production AES. That's not evidence real-world AES is at risk. The report doesn't say whether the same persistence holds when nobody checks in at all.
The more interesting read: capability was apparently never the constraint. Someone had to keep the model on task, and that's a much smaller problem to solve than the cryptography itself. Watch whether Anthropic tries the same experiment with zero human check-ins next time.
Anthropic published the account on July 28.
Each link below shares sources, entities, or timing with this story.
Anthropic's Frontier Red Team published a previously unknown attack on a NIST post-quantum signature candidate, found by a multi-agent system with a human collaborator who is not a lattice cryptography specialist. Claude also produced "Möbius Bridge," a fingerprinting techniqu...
The most capable Claude tier is now government-gated. Let that sink in for a second. Not export-controlled to China, not restricted to enterprise. Gated by Commerce, for a U.S. company's U.S. customers. A June 26 letter from Commerce Secretary Howard Lutnick partially lifted t...
Green's July 29 assessment splits the two results sharply. He calls the HAWK attack significant precisely because HAWK was advancing toward standardization and "none of the ingredients are exotic," meaning the model won by thoroughly applying existing tools rather than inventi...
Anthropic's benchmark with ETH Zurich, Tel Aviv University, University of Haifa, and TU Berlin spans 191 tasks across six primitive families, mostly from AES, SHA-3, Lightweight Cryptography, and Post-Quantum NIST competitions. Tier 1's 49 known-broken schemes saw five frontie...
Anthropic shipped Fable 5 on June 9. Willison spent ~5.5 hours stress-testing it: slow and expensive, but it handled everything he threw at it, including agentic coding. (Simon Willison) The tell that it's a real working model and not a benchmark queen: because it post-dated A...
This one hit my inbox and I had to read it twice. Anthropic announced a partnership with SpaceXAI for the entire Colossus 1 data center in Memphis. 220,000 NVIDIA GPUs. 300+ megawatts. That's the largest single compute acquisition by any AI lab. Full stop. But the part that ma...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.