Tools
vLLM v0.28.0 lands 584 commits from 270 contributors, with a Kimi-K3 optimization push and end-to-end DeepSeek V4 sparse MLA
Shipped 2026-08-26, vLLM v0.28.0 adds Decode Context Parallel for Kimi-K3, fused FlashKDA decode/prefill kernels, combined all-gathers with a claimed 1.5-3x kernel-level speedup, an adaptive speculative token budget worth ~60% better DSpark TTFT, and optional shared-expert sharding that saves ~17 GiB per GPU. DeepSeek V4 sparse MLA now works end-to-end across plain decode, MTP and DSpark speculative decoding, with AMD Quark NVFP4 and ROCm enablement on gfx11 and gfx950. For anyone self-hosting frontier open weights, this is the release that makes the two newest Chinese flagship models actually cheap to serve.
Source
↳ Follow the thread