Reddit
LMSYS Published the Real Qwen3.8-Flash-Next Deployment Numbers: Offloading the 51B N-Gram Table Frees 23.46 GiB Per GPU and Grows KV Cache 78.5%
LMSYS's day-0 SGLang post dated 2026-08-26 gives the numbers the release-day megathread did not: 125B main parameters plus a separate 51.2B Per-Layer Embedding table (about 95.4 GiB in BF16), 6B activated per token, 48 layers split 36 GDN linear attention and 12 QSA sparse attention, 512 experts with top-10 routing. Sparse pinned-host offload of the PLE table on H200 with TP4 drops weight footprint from 83.91 to 60.45 GiB per GPU and raises KV cache capacity from 1.84M to 3.28M tokens. On B200 TP4 with MTP speculative decoding it hits 540 tok/s at batch size 1 with an accept length of 3.3.
Source
↳ Follow the thread