Fetching from the wire…
Models2026-08-08 · source-backed
ARC Prize published verified semi-private results from July 31 testing: 89.0% on ARC-AGI-1 at $0.02/task and 61.4% on ARC-AGI-2 at $0.04/task at max effort, with low effort still reaching 84.0% and 46.0%. Artificial Analysis scored it 50-52 on Intelligence Index and called it the least expensive well-known model to run globally at roughly 3 cents per benchmark test. The number builders should actually read is the gap: cost doubles and accuracy drops 28 points from ARC-AGI-1 to ARC-AGI-2. Cheap high scores on the easier set don't transfer.
Each link below shares sources, entities, or timing with this story.
DeepSeek released deepseek-v4-flash / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (DeepSeek released deepseek-v4-flash); both cover Artificial Analysis, Intelligence Index, July; overlapping topics (cost, task).
Shared entities / Same source domain / Shared topic / Earlier coverage
Both cover AGI, ARC, ARC Prize; reported by the same outlet (arcprize.org); overlapping topics (arc-agi-2, benchmark, doubl).
Shared entities / Same source domain / Earlier coverage
Both cover AGI, ARC, ARC Prize, July; reported by the same outlet (arcprize.org); earlier AGI coverage from 2026-07-12.
Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Both cover AGI, ARC; reported by the same outlet (arcprize.org); overlapping topics (arc-agi-2, benchmark, builder).
Shared entities / Same source domain / Earlier coverage / Tension
Both cover AGI, ARC, ARC Prize; reported by the same outlet (arcprize.org); earlier AGI coverage from 2026-03-12.
Both cover AGI, ARC, ARC Prize; reported by the same outlet (arcprize.org); earlier AGI coverage from 2026-03-10.
Both cover AGI, ARC, ARC Prize; reported by the same outlet (arcprize.org); earlier AGI coverage from 2026-03-07.
Both cover AGI, ARC, ARC Prize; reported by the same outlet (arcprize.org); earlier AGI coverage from 2026-02-12.