Reddit
Speculative Decoding on Qwen3.6-27B Pays Off More at Heavier Quants, Not Less — 10 of 10 Configs Rank the Same Way
An r/LocalLLaMA benchmark writeup (64 upvotes, 18 comments) finishes the speed leg of a speculative-decoding sweep across quantization levels for Qwen3.6-27B and finds a clean monotonic result: the heavier the quant, the more spec-decode buys you, with all ten speculative configurations ranking consistently. That inverts a common assumption that low-bit quants — being memory-bound — benefit most from draft-model acceleration. Single-source and self-reported, but the methodology is stated and the ranking consistency across ten configs is unusually tidy for a hobbyist benchmark.
Source
↳ Follow the thread