Fetching from the wire…
Research2026-09-01 · source-backed
1.5 hours on a single 5090, 44% on ARC-AGI-1 and 7% on ARC-2, 67 cents covering training and inference (mvakde). The method converts puzzle pairs to token sequences, trains autoregressively on evaluation inputs with labels hidden, uses per-task additive embeddings with 3D RoPE and color/dihedral augmentation, then takes the two most common outputs. Every upgrade over the author's previous attempt is a commodity part: SwiGLU, RMSNorm, eight layers instead of four, NorMuon instead of AdamW. Author reports matching TRM and HRM while beating many LLMs.
Each link below shares sources, entities, or timing with this story.
The ARC Prize Foundation is launching ARC-AGI-3, the first major format change since 2019. Unlike versions 1 and 2 (static visual reasoning), version 3 uses game-like environments where agents explore without instructions, discover rules, and adapt to hand-crafted levels that...
ARC-AGI-3 makes a fundamental shift: instead of static puzzle-solving, it measures agency — a model's capacity to set and pursue goals independently in interactive environments. Public release March 25. If frontier models still fail at ARC-AGI-3 despite succeeding at coding ta...
- Source: arcprize.org - Date: 2026-02-10 First interactive reasoning benchmark: agents navigate video-game-like environments with no instructions, discovering rules across 1,000+ levels in 150+ hand-crafted environments. Directly challenges "scale is all you need" by requirin...
ARC Prize published verified semi-private results from July 31 testing: 89.0% on ARC-AGI-1 at $0.02/task and 61.4% on ARC-AGI-2 at $0.04/task at max effort, with low effort still reaching 84.0% and 46.0%. Artificial Analysis scored it 50-52 on Intelligence Index and called it...
Sergey Rodionov's paper tests four Codex-based agent variants to isolate what actually drives performance. Verification (simplification plus exact observation reproduction) ranked highest in every setting, but at substantially higher cost. The textual baseline beat the executa...
OpenAI shipped native computer use in GPT-5.4, scoring 75.0% on OSWorld-Verified vs. the 72.4% human baseline (up from 47.3% in GPT-5.2). This is the first general-purpose model to surpass human performance on real desktop workflows. With 1M token context, 92.8% GPQA Diamond,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.