Fetching from the wire…
Public story · 2026-07-20 · high
The paper's ablation also shows gpt-5.6-sol's near-perfect public scores are memorized test data, not new reasoning.
Why now: The paper arrives as agent builders default to executable world models without testing simpler baselines, and it's the rare study willing to undercut its own headline result.
Sergey Rodionov's ablation on ARC-AGI-3 finds verification beats executable world models in every agent configuration he tested. That's the tradeoff builders now face: across the four Codex-based variants in the paper, verification wins on accuracy. But it costs substantially more to run than the alternatives, so picking a scaffold means weighing performance against compute spend.
Verification here means two things: simplifying the task, and reproducing observations exactly. Isolating that as the top performer in every setting is the paper's main result.
The bigger surprise sits one step down. In some configurations, a plain textual baseline beat the variant built around an executable world model, a scaffold agent designers often reach for by default. Rodionov's ablation cuts against the assumption that a simulated, executable version of the environment is worth the engineering it takes to build.
He also flags gpt-5.6-sol's near-complete score on ARC-AGI-3's public games as test-set saturation, not a capability jump. The model has likely seen those specific puzzles before, so a near-perfect score there says less about reasoning than it looks like it does.
That's what makes the paper's own headline number suspect. A cheaper textual baseline already beats the world-model variant in some settings, so the most engineered approach isn't automatically the best one. If exact-observation-reproduction is really what drives verification's edge, a Codex-based agent should get most of that gain without ever building a world model at all. That's a claim this ablation's own numbers already point toward.
Each link below shares sources, entities, or timing with this story.
OpenAI released Codex / Shared entity: Codex / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenAI released Codex); both cover Codex; overlapping topics (against, agent).
Codex competes with Claude Code / Shared entity: Codex / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Codex competes with Claude Code); both cover Codex; reported by the same outlet (arxiv.org).
Codex competes with Claude Code / Shared entity: Codex / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Codex competes with Claude Code); both cover Codex; overlapping topics (actually, agent).
Cursor benchmarked against Codex / Shared entity: Codex / Shared topic / Earlier coverage
Linked by a graph relationship (Cursor benchmarked against Codex); both cover Codex; overlapping topics (actually, agent, artifact).
JetBrains supports Codex / Shared entity: Codex / Shared topic / Earlier coverage
Linked by a graph relationship (JetBrains supports Codex); both cover Codex; overlapping topics (against, agent, beat).
Codex competes with Claude Code / Shared entity: Codex / Shared topic / Earlier coverage
Linked by a graph relationship (Codex competes with Claude Code); both cover Codex; overlapping topics (agent, alway).
OpenAI released Codex / Shared entity: Codex / Same source domain / Earlier coverage
Linked by a graph relationship (OpenAI released Codex); both cover Codex; reported by the same outlet (arxiv.org).
JetBrains supports Codex / Shared entity: Codex / Shared topic / Earlier coverage
Linked by a graph relationship (JetBrains supports Codex); both cover Codex; overlapping topics (actually, agent).