News
Cloudflare Details How It Serves Kimi K2.6 and GLM 5.2: FP8 KV Cache Doubles Usable Context, INT4 Weights Cut a 705GB Checkpoint to 421GB
In an August 3, 2026 engineering post, Cloudflare published concrete numbers from quantizing frontier open models in production. Dropping the KV cache from BF16 to FP8 (e4m3) raises Kimi K2.6's in-memory context from roughly 686,000 to about 1.37 million tokens and lifts concurrency from 32 to 64 requests, reaching 2,192 tokens/sec — 41% above BF16 peak at roughly 30% less cost per token. Compressing GLM 5.2 weights from FP8 to INT4 shrinks the checkpoint ~40% (705GB → 421GB) and per-GPU memory from ~88GB to ~52GB, with accuracy within 0.8 points across benchmarks and page-tagged cache integrity checks costing under 1% throughput.
↳ Follow the thread