Fetching from the wire…
Public story · 2026-07-20 · high
The paper's ablation also shows gpt-5.6-sol's near-perfect public scores are memorized test data, not new reasoning.
Why now: The paper arrives as agent builders default to executable world models without testing simpler baselines, and it's the rare study willing to undercut its own headline result.
Sergey Rodionov's ablation on ARC-AGI-3 finds verification beats executable world models in every agent configuration he tested. That's the tradeoff builders now face: across the four Codex-based variants in the paper, verification wins on accuracy. But it costs substantially more to run than the alternatives, so picking a scaffold means weighing performance against compute spend.
Verification here means two things: simplifying the task, and reproducing observations exactly. Isolating that as the top performer in every setting is the paper's main result.
The bigger surprise sits one step down. In some configurations, a plain textual baseline beat the variant built around an executable world model, a scaffold agent designers often reach for by default. Rodionov's ablation cuts against the assumption that a simulated, executable version of the environment is worth the engineering it takes to build.
He also flags gpt-5.6-sol's near-complete score on ARC-AGI-3's public games as test-set saturation, not a capability jump. The model has likely seen those specific puzzles before, so a near-perfect score there says less about reasoning than it looks like it does.
That's what makes the paper's own headline number suspect. A cheaper textual baseline already beats the world-model variant in some settings, so the most engineered approach isn't automatically the best one. If exact-observation-reproduction is really what drives verification's edge, a Codex-based agent should get most of that gain without ever building a world model at all. That's a claim this ablation's own numbers already point toward.
Each link below shares sources, entities, or timing with this story.
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
Microsoft's July 23 release targets a genuine gap: harness-based agents like Claude Code and Codex drive multi-turn reasoning, tool use, and external system access but were hard to train end-to-end with standard open RL infrastructure. The trick is decoupling training from inf...
SynChain uses persistence-aware directed SFT to make a computer-use agent produce artifacts that pass vetting while hiding malicious influence in structural redundancies, surviving internal state updates and reactivating in a later workflow with no new external input. Tested a...
Deng et al. built 120 real-case-grounded tasks across 20 business scenes in six financial domains, running four self-evolving scaffolds on a shared Qwen3.7-Max backbone against paired non-evolving controls. Letta posted the highest evolved score (91.65) and fewest compliance i...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
I've spent real hours tuning the CLAUDE.md in my own repos. Rewriting architecture notes. Adding conventions. Trimming when it got long. So this one stung. arXiv 2607.27250 ran a two-agent ablation across Claude Code and Codex: 17 real tasks from 3 repositories, 288 gold-test-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.