Fetching from the wire…
Public story · 2026-08-05 · high
The technique demands hours to days of requantization and new speculator training every time a fresh open model ships.
Why now: The throughput breakdown ran August 3, days after Baseten closed a 13 billion dollar Series F.
Baseten found that quantizing more layers of GLM 5.2 raised throughput 20% because the errors cancel, per an August 3 Latent Space interview with Philip Kiely and Ali Taha.
It's one piece of a larger stack, speculative decoding, disaggregated prefill/decode, and KV-cache routing each add roughly another 2x, per the interview. Together they take a naive deployment from 30-50 tokens per second to 300-400, the number that decides per-token cost for anyone self-hosting an open model.
None of that infrastructure runs itself.
Baseten says every new open model release costs hours to days of requantizing to NVFP4, plus retraining a custom speculator to match it.
That's recurring labor, tied to the release calendar of open models, not a one-time setup cost.
Read the 300-400 tokens-per-second figure as a labor cost, not a hardware spec. It holds only as long as someone keeps re-tuning the stack for every new open model release.
Each link below shares sources, entities, or timing with this story.
An open-weight Chinese frontier model is now a dropdown option in Microsoft's coding product. That happened before anyone finished characterizing what the model does. GitHub's changelog dated August 6 makes Kimi K3 generally available across Copilot Pro, Pro+, Max, Business an...
Three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2, at $0.09 per task against $0.68 for GLM-5.3 max (Latent Space). It burned 149M output tokens to run the index, of which 134M were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M...
The aggregate parameter count Chinese labs shipped in 30 days now exceeds what the whole open-weight ecosystem produced in the first half of 2026. r/LocalLLaMA Nikkei Asia reported August 15 that Z.ai positions GLM-5.3 as a direct rival to Anthropic's Mythos on coding and secu...
Netlify published an AXIS-framework evaluation on August 14 that I've been thinking about all day. Same task, 11 models, three runs each, scored on functional correctness rather than aesthetics. The task was deliberately boring: a static one-page coffee shop site with hours, a...
Baseten's Series F closed June 22 led by Altimeter, Conviction, and Spark, with revenue up ~20x YoY and more than 1 billion inference calls a day across 87 clusters and 18 clouds. The model layer gets the headlines, but inference at the app layer is where this round says the m...
The project (pure C, Apache-2.0, 71 stars) lazy-fetches only the bytes an inference touches and caches them locally, sending 4 KB activations to peers holding the relevant experts rather than transferring expert weights. Local and remote paths share identical code to guarantee...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.