Fetching from the wire…
Public story · 2026-08-05 · high
The technique demands hours to days of requantization and new speculator training every time a fresh open model ships.
Why now: The throughput breakdown ran August 3, days after Baseten closed a 13 billion dollar Series F.
Baseten found that quantizing more layers of GLM 5.2 raised throughput 20% because the errors cancel, per an August 3 Latent Space interview with Philip Kiely and Ali Taha.
It's one piece of a larger stack, speculative decoding, disaggregated prefill/decode, and KV-cache routing each add roughly another 2x, per the interview. Together they take a naive deployment from 30-50 tokens per second to 300-400, the number that decides per-token cost for anyone self-hosting an open model.
None of that infrastructure runs itself.
Baseten says every new open model release costs hours to days of requantizing to NVFP4, plus retraining a custom speculator to match it.
That's recurring labor, tied to the release calendar of open models, not a one-time setup cost.
Read the 300-400 tokens-per-second figure as a labor cost, not a hardware spec. It holds only as long as someone keeps re-tuning the stack for every new open model release.
Each link below shares sources, entities, or timing with this story.
Altimeter invested in Baseten / Shared entity: Baseten / Earlier coverage / Tension
Linked by a graph relationship (Altimeter invested in Baseten); both cover Baseten; earlier Baseten coverage from 2026-06-26.
Baseten supports Kimi K3 / Shared entity: August / Shared topic / Earlier coverage
Linked by a graph relationship (Baseten supports Kimi K3); both cover August; overlapping topics (august, cost, counterintuitive).
Baseten supports Kimi K3 / Shared entity: Latent Space / Same source domain / Earlier coverage
Linked by a graph relationship (Baseten supports Kimi K3); both cover Latent Space; reported by the same outlet (latent.space).
Linked by a graph relationship (Baseten supports Kimi K3); both cover Latent Space; reported by the same outlet (latent.space).
Linked by a graph relationship (Baseten supports Kimi K3); both cover Latent Space; reported by the same outlet (latent.space).
Shared entities / Same source domain / Earlier coverage
Both cover Baseten, GLM; reported by the same outlet (latent.space); earlier Baseten coverage from 2026-07-16.
Both cover GLM, Latent Space; reported by the same outlet (latent.space); earlier GLM coverage from 2026-03-19.
Baseten supports Kimi K3 / Shared entity: Baseten / Earlier coverage
Linked by a graph relationship (Baseten supports Kimi K3); both cover Baseten; earlier Baseten coverage from 2026-07-28.