Tycho Reaches Perfect 100.00 RHAE on ARC-AGI-3 With Opus 5 and GPT-5.6 Sol, Using 61% Fewer Actions Than Humans
Tycho formalizes ARC-AGI-3 games as parameterized rendered deterministic Moore machines and has a coding agent build, test, repair, or bypass a free-form executable hypothesis of each game during play. Across all 25 public games under matched inference budgets, actor-requested delegation to a model builder scored the highest mean Relative Human Action Efficiency at 88.49; with that policy, GPT-5.6 Sol and Opus 5 both hit 100.00 RHAE and completed all 183 levels, with Opus 5 using 61% fewer scored actions than the aggregate official human baselines. Notably, automatic repair after verification failures produced simulators that matched observed transitions far better yet reached only 83.07 RHAE — reproducing dynamics is not the same as identifying the objective.
Source
↳ Follow the thread