Fetching from the wire…
Policy2026-09-13 · source-backed
In an X thread September 12, ARC Prize said ARC-AGI-4 will be "a benchmark for autonomous open-ended innovation," arguing humans still significantly outperform AI at open-ended invention. The same statement takes a side in the week's argument: "Any coordinated effort by the AI industry to reduce openness or concentrate access to frontier AI would undermine that positive-sum future." Nothing about ARC-AGI-4 is on the ARC Prize blog yet, where the newest post is still the September 3 Astra on ARC-AGI-3 writeup, so this is X-only so far.
Each link below shares sources, entities, or timing with this story.
67 on coding against Fable 5.1's 70 in Claude Code. Astra does post a 2% hallucination rate against 9.4% for GPT-5.6 Sol, and 0% scope violations against 48%. Per-task cost runs the other way, $4.72 for Astra against $9.18 for Fable 5.1 at identical $10/$50 list pricing, and A...
An r/OpenAI thread documents an account warning from a user who says they only use Codex for coding, appeal rejected and warning upheld (r/OpenAI). A commenter at 71 upvotes reports the identical email, appealed on the grounds that security assessment is part of app developmen...
The concessions are unusual for him: "we clearly had some missteps as a company. Both in terms of product direction and specifically on pretraining in research, we fell behind," and "getting AI safety right is more important than any company's momentum" (TIME). Concrete items:...
Huang's September 6 post reads "From ChatGPT to o1 to Astra in 4 years. AGI has arrived," noting Astra was trained on more than 100,000 Grace Blackwell NVLink72 systems, and Greg Brockman amplified it saying OpenAI is "now moving into the AGI era" (Business Insider). Marcus re...
He conceded real ground on September 3, calling it vindicating "to see that a product from OpenAI explicitly creates and manipulate symbolic world models" after a decade arguing for neurosymbolic approaches. He then attacks the rollout on two fronts: the system is less monitor...
The 48-level clear took r/OpenAI's top slot at 1,124 upvotes; a 166-upvote r/singularity post put Astra at 13% on MazeBench without tools, and a smaller thread reported over-engineering problems in Unity. The spread is the useful read: strong on tool-mediated multi-step browse...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.