Fetching from the wire…
Top 5 · 2026-08-02 · source-backed
The number that reframes everything isn't ten. It's two thousand.
OpenAI published "Ten advances in mathematics and theoretical computer science" on August 1, claiming an internal version of Astra produced new results on ten problems that had seen no progress on the main result for at least a decade. High-dimensional sphere packing. Connes's rigidity conjecture. Ehrhart's volume conjecture. Multicolor Ramsey numbers. Quantum parallel repetition. Arithmetic circuit complexity. The closest vector problem. Non-sofic groups. Binary and spherical codes. Extremal number conjectures.
They shipped receipts. The openai/ten-proofs repo is Apache-2.0, Lean 4.32.0, sitting around 275 stars, with machine-checkable certificates alongside an LLM-written PDF reconstructing the derivations from unpublished reasoning traces. Machine-checkable is the part that separates this from every previous "AI solved math" claim. You don't have to trust the narrative. You can run the kernel.
Then Noam Brown posted the cost: all ten proofs, combined, ran under $2,000 at GPT-5.6 Sol API prices. He also killed the obvious follow-up question directly. "Sadly no Millennium Prize problems (yet)." They tried other major problems and failed. They didn't spend heavily per problem, and test-time compute could be pushed much further.
That last clause is the actual story. If frontier mathematical discovery costs $200 a problem and nobody has pushed the compute knob hard, then this isn't a capability that arrives with the next model. It's a dial that's already installed and turned down. That's a completely different planning assumption than "wait for GPT-6."
The mathematicians are not celebrating. Terence Tao described GPT-5.6 Pro solving problems he'd personally spent time on as "very strange and not particularly pleasant." Timothy Gowers, a Fields medalist, warned about the "possible destruction of mathematical culture" as the literature expands while human understanding thins. Queen Mary's Abhishek Saha landed somewhere more useful: frontier models are "at least as good as a solid and indefatigable PhD student," and he now plays "conductor, rather than doubling up as the whole orchestra." Kirwin Hampshire called the reassuring framing "a well-muffled scream," and his July essay resurfaced on r/singularity today at 356 upvotes and 513 comments, a 1.44 comment-to-score ratio that means people are arguing, not nodding.
Willison flagged the calibration problem nobody else did: OpenAI discloses cost per success but not the denominator. How many problems did Astra attempt? How many did it fail? Ten wins out of ten attempts and ten wins out of four hundred are the same press release and completely different technologies.
I'd add a second gap. Thomas Bloom at Manchester called it "big news" while noting nobody has had time to referee these at the depth these conjectures normally get. A Lean certificate proves the theorem follows from the axioms. It doesn't tell you how much human problem-shaping happened before the run, which is exactly what Tao's "tireless literature-scanning assistant" framing is pointing at.
What to do with this: stop treating frontier reasoning as a fixed capability tier you shop for. If the cost curve on hard-problem solving is this low and this unexplored, the question for your own work isn't "which model" but "how much compute am I willing to spend on one problem." Most of us have never asked. And when a rebuttal PDF claiming the Connes disproof was invalid got shredded on HN within hours, the same author has a claimed Riemann Hypothesis proof, and commenters flagged the rebuttal itself as likely AI-generated, the real open question surfaced: when validation costs this much effort, how does anything get adjudicated at all?
Each link below shares sources, entities, or timing with this story.
Hugging Face criticizes OpenAI / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Hugging Face criticizes OpenAI); both cover GPT, July, LLM, Most; reported by the same outlet (simonwillison.net).
OpenAI released Codex / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (OpenAI released Codex); both cover GPT, July, LLM, OpenAI; earlier GPT coverage from 2026-07-26.
Hugging Face criticizes OpenAI / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Hugging Face criticizes OpenAI); both cover July, OpenAI, Willison; reported by the same outlet (simonwillison.net).
Copilot uses GPT / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Copilot uses GPT); both cover GPT, July, OpenAI, Under; overlapping topics (frontier, model, problem).
OpenAI released Luna / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (OpenAI released Luna); both cover GPT, High, July, OpenAI; overlapping topics (capability, frontier, openai).
LLM uses OpenAI / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover LLM, OpenAI, Willison; reported by the same outlet (simonwillison.net).
Sam Altman works at OpenAI / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Sam Altman works at OpenAI); both cover GPT, High, OpenAI; reported by the same outlet (x.com).
OpenAI uses Claude Code / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (OpenAI uses Claude Code); both cover August, GPT, OpenAI, Willison; reported by the same outlet (simonwillison.net).