Reddit
DeepSeek V4 Flash 0731 Posts 89.0% on ARC-AGI-1 at 2 Cents a Task — Verified Semi-Private Scores Put It on the Cost-Efficiency Frontier
ARC Prize published verified semi-private results for DeepSeek V4 Flash 0731, tested July 31, 2026: 89.0% on ARC-AGI-1 at $0.02 per task and 61.4% on ARC-AGI-2 at $0.04 per task at max effort, with three reasoning variants scored (low effort still hits 84.0% / 46.0%). Artificial Analysis separately scored the model 50-52 on its Intelligence Index and called it the least expensive well-known model to run globally at roughly 3 cents per benchmark test. For builders the takeaway is the ARC-AGI-2 gap: cost doubles and accuracy falls 28 points versus ARC-AGI-1, so cheap high scores do not transfer to the harder set.
↳ Follow the thread