Tools
LMDeploy 0.17.0 integrates DeepEPv2 and adds Kimi K2.6 plus a Mooncake store KV connector
Published 2026-09-01, LMDeploy v0.17.0 integrates DeepEPv2 (#4783), adds PyTorch-engine support for Kimi K2.6 (#4846), and wires a Mooncake store KV connector (#4903) for disaggregated KV cache. Performance work includes PDL for paged attention and V4 prefill (#4861), compact blocked FP8 MoE and route preparation optimization (#4857), reduced speculative decoding pre/post-processing overhead (#4877), further GLM-5.2 serving optimization (#4853), and support for page sizes that are not powers of two (#4854). Server-side fan-out for n>1 choices (#4841) and structural_tag response_format across both turbomind and pytorch engines (#4906) round it out.
Source
↳ Follow the thread