Reddit
Epoch's FrontierMath Erdős benchmark: GPT-6 Astra solves 2 of 68 open problems, every other model scores zero
Epoch AI launched FrontierMath Erdős on September 1, built from 68 unsolved Erdős problems that Thomas Bloom picked out of roughly 652 in his catalog, with 50 already formalized in Lean via Google's Formal Conjectures and 18 newly formalized. Each attempt gets a fixed budget of $300 and 72 hours. GPT-6 Astra scored 3% by disproving one problem and proving another; GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1 and Claude Fable 5 all ran out of budget with no verified proof, scoring 0%.
Source
↳ Follow the thread