Reddit
Opus 5 Scores 30.2% on ARC-AGI 3, Nearly 4x the Previous Best of 7.8% and 20x Opus 4.8's 1.5%
Claude Opus 5 posted 30.2% on ARC-AGI 3, against a prior field best of 7.8% held by GPT-5.6 Sol at max reasoning effort and just 1.5% from Anthropic's own Opus 4.8. François Chollet designed ARC-AGI explicitly to resist the memorization and benchmark-targeting that inflate scores elsewhere, which makes a jump of this size on a single generation unusual rather than routine. The ARC Prize leaderboard hit the Hacker News front page within hours of the model's release (106 points), and the result drove the top posts on r/singularity and r/ClaudeAI. Caveat worth holding: a 30.2% score still means the model fails roughly seven of ten tasks a human solves easily.
↳ Follow the thread