Fetching from the wire…
Public story · 2026-09-18 · high
Vercel's August data shows open-weight models crossed a majority of gateway tokens, up from 7% in December, while Anthropic keeps 64% of the spending.
Why now: Vercel published the numbers on September 17, covering August traffic.
Open-weight models crossed a majority of AI Gateway token volume in August, reaching 56%, per Vercel's September AI Gateway Production Index. Nine months earlier, in December 2025, that share was 7%. For any team defaulting high-volume calls to a frontier model, that gap is real budget sitting on the table.
The money didn't follow the tokens. Open weights accounted for just 14% of dollars spent on the gateway. Anthropic alone took 64% of all spending, with Claude Opus 5 responsible for 22.5% by itself.
Average token cost fell 23.2% in August alone and more than 50% over five months.
In matched twelve-day windows after launch, GPT-6 Astra took 7.7% of gateway spend against Fable 5.1's 3.7%, even though twice as many teams used Fable.
Read together, the volume and spending numbers describe two tiers splitting further apart. Open weights are absorbing classification, extraction, routing, and summarization, the work you run thousands of times a day where cost matters more than quality. Frontier closed models keep the calls where a wrong answer is expensive.
This is first-party telemetry from a gateway that sees real production traffic, which makes it a cleaner read than a survey. It's also Vercel's own customer base, skewed toward companies that already deploy there. Treat the absolute percentages as directional and the trend line as the real signal.
The shift shows up in what's getting built for that cheap tier. Cactus Compute's Needle 3 is a 121M-parameter tool-calling model that ships in 8-29MB slices. It returns a filled-in function call or a typed record, and returns an empty list rather than guessing when no declared tool fits.
It's a component built to return a type, not a chat model that returns a string, a different engineering problem than prompt tuning.
Each link below shares sources, entities, or timing with this story.
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
Per-token prices went down at both labs. Subscriptions are draining faster at both labs. Those aren't in tension once you look at token counts. On the OpenAI side, r/OpenAI collected reports from Linux.do and NodeSeek alleging Astra consumes more Plus quota than its published...
Two facts sit next to each other and neither cancels the other out. Anthropic published on September 4 that an internal general-purpose research model, roughly comparable to Claude Fable 5.1, formalized Fermat's Last Theorem in Lean over 11 days working largely autonomously. T...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.