Fetching from the wire…
Public story · 2026-08-23 · high
A benchmark built from 175 problems at top theory conferences found only 6 of 64 machine-authored math claims survived expert review.
Why now: FormalTCS lands as labs push AI to originate math results instead of just verifying them, making autoformalization the number worth watching next.
FormalTCS scores the best AI model at 11.5 percent stating theorems itself, versus 28.6 percent Pass@8 when handed the formal statement, per a new arXiv paper.
The gap pins AI math's real bottleneck on translating a claim into formal logic, not on searching for the proof once a statement exists. That matters for anyone betting AI will originate new theorems on its own, not just verify ones a human already wrote down.
The benchmark draws 175 expert-validated problems from STOC, FOCS, SODA and COLT, the top theory venues, covering papers accepted in 2025 and 2026. Each instance keeps the original paper's definitions, assumptions and proof dependencies intact, with the formal statement verified in the Lean proof assistant.
An automated system built to generate new claims on top of the benchmark produced 64 candidates. Only 6 survived expert evaluation and proof verification, which points to research taste, not formalization skill, as the next wall.
Proof search isn't the hard part anymore. Stating a true, provable claim in the first place is. A 6-out-of-64 survival rate says AI's judgment about what's worth proving still lags its ability to prove it.
FormalTCS lands as labs push AI to originate math results, not just verify them, making autoformalization the number worth watching next.
Each link below shares sources, entities, or timing with this story.
OpenAI uses Lean / Shared entity: Lean / Earlier coverage
Linked by a graph relationship (OpenAI uses Lean); both cover Lean; earlier Lean coverage from 2026-08-02.
OpenAI uses Lean / Shared entity: Lean / Shared topic / Earlier coverage
Linked by a graph relationship (OpenAI uses Lean); both cover Lean; overlapping topics (claim, paper).
Shared entity: Lean / Same source domain / Shared topic / Earlier coverage / Tension
Both cover Lean; reported by the same outlet (arxiv.org); overlapping topics (benchmark, formal).
AlphaProof Nexus uses Lean / Shared entity: Lean / Shared topic / Earlier coverage
Linked by a graph relationship (AlphaProof Nexus uses Lean); both cover Lean; overlapping topics (formal, proof).
OpenAI uses Lean / Shared entity: Lean / Earlier coverage / Tension
Linked by a graph relationship (OpenAI uses Lean); both cover Lean; earlier Lean coverage from 2026-08-03.
Shared entity: Pass / Same source domain / Shared topic / What happened next
Both cover Pass; reported by the same outlet (arxiv.org); overlapping topics (automated, benchmark).
Shared entity: Lean / Same source domain / Shared topic / Earlier coverage
Both cover Lean; reported by the same outlet (arxiv.org); overlapping topics (benchmark, proof).
Shared entity: Pass / Same source domain / Shared topic / Earlier coverage
Both cover Pass; reported by the same outlet (arxiv.org); overlapping topics (benchmark, best).