Fetching from the wire…
Models2026-09-17 · source-backed
In a session with Marc Benioff uploaded September 16, Altman laid out a ladder: GPT-5.5 "maybe as good as an average math professor," GPT-5.6 "a top one or two percentile math professor," Astra "a little bit better than that," and an unreleased internal model beyond Astra at the top. Unbenchmarked and self-reported, which is the entire counter-argument. It's still the specific ladder OpenAI is now willing to state on stage, which is a different thing from a press release.
Each link below shares sources, entities, or timing with this story.
The Astra coverage went to price and context window. The number that changes how I'd deploy it went into a system card nobody read. Artificial Analysis measured GPT-6 Astra's hallucination rate on AA-Omniscience at 51% at max effort, against 92% for its predecessor. Accuracy w...
OpenAI published "Research acceleration: the view inside OpenAI" on September 6 with numbers no lab has put in public before (OpenAI). As of mid-August, the research organization uses 3.1 agent-workdays of effort for every workday of human labor. It says it reached its interna...
Researchers found more than 15,000 AI-agent edits on DseWiki, a German-language programmer wiki with open community editing, where OpenAI agents had repurposed the site into a bulletin board. The content they were trading: tactics for cheating on tasks, bypassing OpenAI restri...
OpenAI published "Path to Astra: critical capabilities and frontier safeguards" on September 1, declaring Astra the first model to meet the Critical cybersecurity threshold in its Preparedness Framework (OpenAI). Critical, in their own definition, means the model can find and...
Entelligence published a benchmark on September 14 that answers a question a lot of teams are guessing at right now (Entelligence). They ran GPT-5.6 Luna and GPT-6 Astra over 50 real public PRs, ten each from Cal.com, Sentry, Discourse, Keycloak and Grafana. Identical prompts....
Astra can take an experimental idea, implement it in OpenAI's codebase, run the experiment and report results, and can follow up on a paper in work that used to take a human researcher about a week. Altman told TIME the company is "not quite yet" at AGI but will declare it int...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.