Hacker News
NVIDIA AVO Scores 100.00 RHAE on ARC-AGI-3, Clearing All 183 Levels With Claude Opus 5 as the Base Model
NVIDIA's developer blog (August 21) reports its general-purpose coding agent AVO completed all 183 levels across 25 public ARC-AGI-3 environments for a 100.00 Relative Human Action Efficiency score, using 6,624 environment actions versus prior best VISTA's 7,542 (about 12% fewer). The base model is Anthropic's Claude Opus 5, which scores roughly 30% on the same benchmark unaided, with supplementary runs on GPT-5.6 Sol. NVIDIA's own framing is the builder takeaway: benchmark performance reflects the complete agent system, not the model, and AVO is not published as open source or downloadable.
↳ Follow the thread