Fetching from the wire…
Top 5 · 2026-03-29 · source-backed
GPT-5.4 scores 0.26%. Opus 4.6 scores 0.25%. Grok-4.20 scores 0.00%. Humans score 100%. The Decoder covered the ARC-AGI-3 launch on March 25, and the results make every "AGI is here" claim look premature.
François Chollet launched ARC-AGI-3 at Y Combinator HQ alongside a fireside conversation with Sam Altman. The benchmark is fully interactive: hundreds of game-style environments with no instructions, no rules, and no stated goals. Agents must figure out what to do by exploring and learning from environmental feedback. Sustained sequential reasoning, state tracking across hundreds of steps, real-time adaptation. Everything current language models can't do.
Here's what stopped me cold: simple CNN and graph-search approaches scored 12.58%. Over 30x better than any frontier LLM. Not a fine-tuned model, not a billion-dollar training run. Basic pattern matching over a narrow domain demolishes trillion-parameter models on tasks requiring actual novel reasoning. The models interpolate beautifully from training data. They can't extrapolate at all. The gap between "looks smart" and "is smart" has never been measured this precisely.
The Chollet-Altman pairing is notable. They agreed on a timeline: AGI "probably by early 2030s, around ARC-AGI 6 or 7" (OfficeChai). The creator of the hardest AI benchmark and the CEO of the company most invested in scaling sat together and acknowledged that current approaches aren't sufficient. Meanwhile, GPT-5.4 saturated USAMO 2026 at 95%, up from near-zero last year (MathArena). So models are getting dramatically better at pattern-matchable math competitions while scoring essentially zero on tasks requiring genuine novel reasoning. That divergence is the whole story.
ARC Prize 2026 offers $2M+ in prizes. All solutions must be open-sourced. If you want to work on the hardest unsolved problem in AI, the benchmark and the funding are waiting. This connects directly to the Anthropic harness story: if individual models can't reason sequentially, you build systems where reasoning is distributed across specialized agents with external feedback loops. Multi-agent architectures aren't just a design pattern. They're a workaround for a fundamental limitation.
Each link below shares sources, entities, or timing with this story.
Opus built by Anthropic / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Opus built by Anthropic); both cover AGI, ARC, GPT, March; overlapping topics (agent, arc-agi-3, benchmark, model).
Opus built by Anthropic / Shared entities / Same source domain / Shared topic / What happened next / Tension
Linked by a graph relationship (Opus built by Anthropic); both cover Anthropic, GPT, Opus; reported by the same outlet (officechai.com).
LLM uses OpenAI / Shared entities / Shared topic / What happened next / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover Anthropic, GPT, LLM, Opus; overlapping topics (benchmark, gpt-5, hardest, model).
Opus built by Anthropic / Shared entities / Shared topic / What happened next
Linked by a graph relationship (Opus built by Anthropic); both cover Anthropic, GPT, March, Multi; overlapping topics (agent, model).
Claude Code uses Opus / Shared entities / Same source domain / Shared topic / What happened next
Linked by a graph relationship (Claude Code uses Opus); both cover AGI, ARC, GPT, Opus; reported by the same outlet (officechai.com).
Simon Willison released LLM / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Simon Willison released LLM); both cover AGI, Altman, Anthropic, ARC; earlier AGI coverage from 2026-02-12.
Sam Altman supports Anthropic / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Sam Altman supports Anthropic); both cover Altman, GPT, Opus, Sam Altman; overlapping topics (gpt-5, model).
LLM uses OpenAI / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (LLM uses OpenAI); both cover AGI, Anthropic, ARC, LLM; overlapping topics (feedback, model).