vLLM Ships Day-0 Kimi K3 Support: 118 tok/s Baseline, 370 tok/s With DSpark Speculative Decoding on 16x GB300 NVL72
The vLLM team published day-0 K3 support on July 27 alongside the v0.26.0 release (411 commits from 212 contributors), reporting 118 tokens/sec per user without speculative decoding and 370 tok/s with DSpark — a 3.14x improvement — on 16 NVIDIA GB300 NVL72 GPUs. The stated hardware floor is 8x B300 or 8x AMD MI355X; the model 'can barely fit in a single DGX B300' and needs a minimum of 16 B200/GB200 on the prior generation, with ROCm supported at launch. Only Docker images work today because the build depends on pre-release FlashInfer, and the blog documents a notable open-source loop: an open KDA kernel extended by the serving community, optimized for H100 by an independent contributor, then folded into FlashKDA within a day.
↳ Follow the thread