ExploitBench: Claude Mythos Achieves Arbitrary Code Execution on 21 of 41 Real V8 Vulnerabilities — GPT-5.5 Manages Only 2
The Decoder / arXiv·high signal
A new arXiv paper (2605.14153) introduces ExploitBench, a 16-flag capability ladder benchmark testing AI models on full exploit synthesis against 41 real V8 JavaScript engine vulnerabilities. Claude Mythos Preview scored 9.90/16 and reached arbitrary code execution on 21 of 41 bugs; GPT-5.5 scored 5.51/16 hitting the top tier on only 2. Researcher Seunghyun Lee noted Mythos 'developed an exploit technique that Lee and a colleague had previously dismissed as too complex.'