Reddit
llama.cpp Merged DFlash2 Speculative Decoding, Reporting 1.81x on an M5 Pro With No Extra Flags
PR 27342 by SubSir and Jian Chen, merged August 27, adds grouped dynamic depthwise convolution and a candidate selector for DFlash2 draft models. Reported numbers are a 1.81x speedup on Apple M5 Pro running Qwen3.8-27B Q4_K_M at roughly 5.03 token acceptance, 1.77x to 1.85x across BF16 and Q8_0 variants, and one Nvidia user reporting a consistent 2x decode speedup at every depth measured, holding at 32k context. The implementation auto-enables for DFlash2 checkpoints, so builders get it without touching their launch command.
↳ Follow the thread