Fetching from the wire…
Public story · 2026-09-17 · high
He framed it as a four-tier ladder, from an average math professor up through Astra, on stage with Marc Benioff on September 16.
Why now: Altman made the comparison during a Dreamforce conversation with Benioff on September 16, and it was still moving through r/singularity a day later.
Sam Altman ranked OpenAI's models against human mathematicians, topping out at an unreleased model past Astra, in a Dreamforce conversation with Marc Benioff.
The claim matters because a separate forecast on r/singularity gives AI just a 50% chance of solving a Millennium Prize Problem by 2054. Altman's ladder, if it holds, would clear that bar decades early.
Altman walked Benioff through four models on a math-ability ladder: GPT-5.5, GPT-5.6, Astra, and one more, unreleased, beyond Astra. GPT-5.5, he said, is "maybe as good as an average math professor." GPT-5.6 reaches the top one or two percentile. Astra tops that. And the unreleased model, Altman said, "can do things that the best mathematicians in the world cannot."
None of it comes with a benchmark, a paper, or a number. It's Altman's own framing, made during a Dreamforce conversation with Marc Benioff that Salesforce posted on September 16. A company ranking its own unreleased model above the best living mathematicians is a promise, not a proof. No outside mathematician has checked it.
OpenAI hasn't said when, or whether, the model beyond Astra becomes available to test.
Each link below shares sources, entities, or timing with this story.
In a session with Marc Benioff uploaded September 16, Altman laid out a ladder: GPT-5.5 "maybe as good as an average math professor," GPT-5.6 "a top one or two percentile math professor," Astra "a little bit better than that," and an unreleased internal model beyond Astra at t...
Fireship's September 1 video resolves the Ox Alpha mystery, priced roughly 40x below Claude (video). The volume figure is what stopped me: a large share of agentic token traffic silently rerouted to a Chinese open-weights model on price alone, before anyone knew what it was. W...
One-shot a full Node.js and React application through DSH pointed at V4-Pro on max settings. Fireship He describes the architecture as everything-is-a-plugin: model adapter, tools, sandbox, UI, and the agent's central while loop are all swappable packages configured in one lin...
Matthew Berman relays a workflow the Cursor team described to him: a separate bot per project and per workstream, tasks triggered from Slack, and the bot invoking the Cursor Agent CLI and cloud agents rather than driving Cursor directly. Their stated reason is context persiste...
Justin Wang and Dan Robinson's RSI Simulator is a browser game where you run an AI lab allocating labor, compute and data toward superintelligence, built on the Elasticity Institute's economics of recursive self-improvement (Paradigm). The model turns on elasticities, chiefly...
7-minute field report published August 11, framing deflationary: the gap between demo reels and lab reality is wider than the funding narrative implies (Fireship). 616,000 views in roughly 24 hours, making it the most-watched skeptical take on embodied AI this week and a usefu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.