Fetching from the wire…
Research2026-09-19 · source-backed
Epoch AI logged the first Major Advance tier solution, on emptiness of the core in approval-based committee elections, open since Aziz, Brill and colleagues posed it in 2017. Becker, Greger and Peters worked interactively with GPT-6 Astra, and Peters said he doubts the team would have found the proof without it. Epoch classifies it as human plus AI. Epoch AI The proof shows an empty core is impossible, so the benchmark problem as stated had nothing to find.
Each link below shares sources, entities, or timing with this story.
Two facts sit next to each other and neither cancels the other out. Anthropic published on September 4 that an internal general-purpose research model, roughly comparable to Claude Fable 5.1, formalized Fermat's Last Theorem in Lean over 11 days working largely autonomously. T...
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
The 48-level clear took r/OpenAI's top slot at 1,124 upvotes; a 166-upvote r/singularity post put Astra at 13% on MazeBench without tools, and a smaller thread reported over-engineering problems in Unity. The spread is the useful read: strong on tool-mediated multi-step browse...
OpenAI conceded its prior disclosures were "ad hoc and less frequent than ideal" and launched a Model Misalignment Reporting Framework, sorting cases into Ready for Disclosure, Minor Investigation, and a Slow Track for complex third-party work. Six incidents from the last six...
The MIT-licensed model-gateway plugin routes GPT requests to OpenAI on the user's ChatGPT login and everything else to Anthropic on the normal claude.ai login, so GPT models appear in /model next to Opus and Sonnet with no API keys. The author has run Astra as the main orchest...
The Astra coverage went to price and context window. The number that changes how I'd deploy it went into a system card nobody read. Artificial Analysis measured GPT-6 Astra's hallucination rate on AA-Omniscience at 51% at max effort, against 92% for its predecessor. Accuracy w...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.