Hacker News
DFlash 2 Adds a Path Selector and a Two-Tap Convolution to Speculative Decoding, Reaching 3.4x Throughput on Qwen3.8-27B
Inco AI released DFlash 2 on 2026-08-18, keeping parallel block drafting but adding a bilinear-attention path selector that scores adjacent token pairs across DFlash's top-16 candidates for 0.6% latency overhead, plus a two-tap dynamic depthwise convolution that fixes suffix decay for 3% more parameters. The result is a 21% gain in mean acceptance length over DFlash, 2.7-3.4x throughput versus autoregressive decoding on Qwen3.8-27B, and 3.1-4.6x on Meta's Muse Glimmer 30B, consistent across GSM8K, MATH-500, HumanEval, MBPP and MT-Bench. Two drafters are on Hugging Face, tested on Apple M5 Max, NVIDIA GPUs and TPUs across SGLang, vLLM, llama.cpp, Ollama and oMLX.
↳ Follow the thread