Windowed-MTP: Sliding-Window Draft Attention Cuts Speculative Decoding Cost 28-44% at 1M-Token Context
Alagappan Valliappan shows that built-in Multi-Token-Prediction draft heads run full attention over the entire KV cache at every draft step, so at million-token context the 'negligibly cheap' draft dominates cost and can make deep native drafts net-negative. Applying a StreamingLLM-style sliding window plus attention sink to the draft's attention only — training-free, drop-in, lossless because the full-attention target still verifies every token — drops ~99% of draft KV entries at 1M and cuts per-decode-step cost 28% to 44% across Qwen GDN-MoE 35B/122B and a Mamba2-hybrid NoPE 120B in SGLang on a single GPU. The unread draft KV (7.7-11% of total KV at 1M) is reclaimed via a ring buffer at no acceptance or quality cost.
↳ Follow the thread