Sources
Baseten's Inference Masterclass: Quantizing MORE Layers Raised GLM 5.2 Throughput 20% Because the Errors Cancel
Latent Space published an inference-engineering deep dive with Philip Kiely and Ali Taha of Baseten on August 3, 2026, days after the company's $13B Series F. The counterintuitive headline: quantizing more layers rather than fewer can improve output quality because the introduced errors cancel each other out — Baseten reports 20% higher throughput on GLM 5.2 with quality held. They stack roughly 2x per technique across speculative decoding, disaggregated prefill/decode, and KV-cache routing to move a naive 30–50 tok/s deployment to 300–400 tok/s, and describe the hours-to-days of requantization (to NVFP4) and custom speculator training required every time a new open model drops.
Source
↳ Follow the thread