Fetching from the wire…
Public story · 2026-09-02 · high
OpenRouter's anonymous "Ox Alpha" traffic spike was Zhipu's GLM-5.3-Flash, priced roughly 40x below Claude.
Why now: Fireship's video, posted September 1, resolved a question that had been sitting open on OpenRouter.
OpenRouter ran an anonymous model under the label "Ox Alpha," and it pulled in 42 trillion tokens over six days. A large share of the platform's agentic traffic rerouted to that unnamed model, on price alone, before anyone knew what it was.
Fireship's video identifying GLM-5.3-Flash, posted September 1, names Zhipu as the maker. The model is Chinese, open-weights, and priced 40 times cheaper than Claude.
For teams building agent pipelines that call out to LLM APIs, the routing here wasn't loyal to a brand. It picked the cheapest option that cleared the bar, and the volume moved before most users knew the name behind it.
What the video doesn't settle is whether that traffic stays on GLM-5.3-Flash now that the name is public. Some of it may have been people testing an unlabeled option out of curiosity rather than production workloads that persist. Fireship doesn't break the volume down by use case, so there's no way to tell exploration from a durable shift in spend.
Whatever preference people claim to have for a given model, 42 trillion tokens of real usage didn't check the label first.
Each link below shares sources, entities, or timing with this story.
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
For five days the biggest launch in OpenRouter's history had no author. Ox Alpha appeared as an uncredited listing over the weekend, more than doubled DeepSeek's usage on the platform, and sent a small army of people running compression-distance tests and tokenizer probes to f...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
OpenAI cut GPT-5.6 Luna roughly 80%, from $1 to $0.20 per million input and $6 to $1.20 output. Anthropic priced Opus 5 at $5/$25 per million, half of Fable 5. The trigger is DeepSeek, Zhipu's GLM-5.2 and Moonshot's Kimi K3 landing 60-90% below US flagship pricing, with DoorDa...
Nathan Lambert doesn't hand out "step change" lightly, so when his June 22 Interconnects essay called GLM-5.2 "the step change for open agents," I read it twice. His argument is sharper than the usual "strong open model" take. Static intelligence benchmarks stopped mattering m...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.