Fetching from the wire…
Models2026-08-08 · source-backed
ARC Prize published verified semi-private results from July 31 testing: 89.0% on ARC-AGI-1 at $0.02/task and 61.4% on ARC-AGI-2 at $0.04/task at max effort, with low effort still reaching 84.0% and 46.0%. Artificial Analysis scored it 50-52 on Intelligence Index and called it the least expensive well-known model to run globally at roughly 3 cents per benchmark test. The number builders should actually read is the gap: cost doubles and accuracy drops 28 points from ARC-AGI-1 to ARC-AGI-2. Cheap high scores on the easier set don't transfer.
Each link below shares sources, entities, or timing with this story.
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
The biggest single-benchmark jump in a frontier model update: 37.6% → 68.8% on ARC-AGI-2. For comparison, GPT-5.2 scored 54.2%, Gemini 3 Pro 45.1%. ARC-AGI-2 measures novel problem-solving on adversarially constructed tasks. A near-doubling suggests genuine reasoning improveme...
While the capability stories pile up, here's the counterweight. As SWE-bench Verified scores cluster near saturation on July leaderboards, an enhanced analysis (SWE-Bench+, on the AIware 2026 benchmark track) found 60.83% of commonly resolved issues contain solution leakage ri...
Built on a post-transformer BDH architecture that reasons recurrently in latent space, it hit 29.5% pass@2 on public ARC-AGI-1 at a computed cost roughly 11x cheaper per task than GPT-5.6 Luna Low, even after OpenAI's 80% price cut on July 30 (Pathway). Be precise about what t...
Three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2, at $0.09 per task against $0.68 for GLM-5.3 max (Latent Space). It burned 149M output tokens to run the index, of which 134M were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.