Fetching from the wire…
Policy2026-09-19 · source-backed
His September 18 guest post makes the case that mathematics should give academic credit to motivated explanations, not only to proof generation: "When proofs can be generated without that understanding, it undermines their value as a proxy." His worked example is Liam Price's solution to Erdős Problem 1196, which came out of Price's interaction with GPT-5.4 Pro but only became useful once researchers interpreted the approach and cleaned the proof into human-readable form. Terence Tao's blog He concedes exposition has no verifier: "There will never be Lean for motivated explanations."
Each link below shares sources, entities, or timing with this story.
Quanta's August 3 piece tallies the assault: OpenAI found a counterexample to Erdős's 1946 unit distance conjecture on May 20, then Astra produced 10 further advances. Google DeepMind evaluated 700 open conjectures in January, solving four and recovering nine forgotten solutio...
Two facts sit next to each other and neither cancels the other out. Anthropic published on September 4 that an internal general-purpose research model, roughly comparable to Claude Fable 5.1, formalized Fermat's Last Theorem in Lean over 11 days working largely autonomously. T...
The number that reframes everything isn't ten. It's two thousand. OpenAI published "Ten advances in mathematics and theoretical computer science" on August 1, claiming an internal version of Astra produced new results on ten problems that had seen no progress on the main resul...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
This one rearranged my week. An essay published August 4 walks through Databricks' independent benchmark of coding harnesses against its own multi-million-line codebase. Pi, a harness with four built-in tools and a system prompt under 1,000 tokens, paired with Opus 4.8 at xhig...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.