Qwen3.8-Flash-Next multi-token prediction merged into ik_llama.cpp, roughly doubling decode on a 5090
r/LocalLLaMA (162 upvotes)·low signal
The merge supports both an integrated MTP head and a separate -md draft file, with the poster reporting 45 to 90 tok/s on a 5090 with 128GB system RAM and working configurations down to a 12GB 4070. A separate Part 2 post on the 2x3090 plus DDR4 build reports 37-41 t/s decode using UD-Q4_K_XL plus expert cache plus MTP, up from the 25-29 t/s that same builder reported two days ago. Both are single-builder measurements on hand-built branches, not vendor numbers.