Sources
ARC-AGI-3 Ablation: Verification Beats Executable World Models, and a Plain Text Baseline Sometimes Beats Both
Sergey Rodionov's July 16 paper (arXiv 2607.15439) tests four Codex-based agent variants on ARC-AGI-3 to isolate what actually drives performance. The verification treatment — simplification plus exact observation reproduction — ranked highest in every setting but at substantially higher cost, while the textual baseline beat the executable-world-model variant in some configurations, undercutting the assumption that executable artifacts are always the right agent scaffold. The author explicitly flags that gpt-5.6-sol's near-complete public-game performance represents test-set saturation, not new capability.
Source
↳ Follow the thread