Fetching from the wire…
Public story · 2026-07-27 · high
Ten speculative-decoding setups ranked the same way across every quant level, on a single self-reported Reddit test.
Why now: The thread was circulating on r/LocalLLaMA as of July 27, 2026.
Speculative decoding speeds up heavier-precision quants more than lighter ones, a sweep of ten configs on Qwen3.6-27B found, per a post on r/LocalLLaMA. Local model builders default to the lightest quant they can run for speed. This test says that instinct works against you once a draft model enters the picture.
The sweep ran ten speculative-decoding setups across multiple quantization levels and found the same ranking held every time: the heavier the quant, the bigger the speedup from the draft model. That's a monotonic result, unusually consistent for a single hobbyist test.
The common wisdom runs the other way. Low-bit quants are the ones assumed to be memory-bound, and memory-bound decoding is exactly what speculative decoding is supposed to fix, so the smallest models were expected to gain the most. This benchmark flips that.
It's one Reddit thread, self-reported, with no independent replication. The methodology is stated and the ranking holds across all ten configs, which is more rigor than most hobbyist benchmarks show, but it's still a single source.
If the ranking holds up outside this one thread, running the smallest quant for speed is the wrong call once speculative decoding is in the mix. Worth watching whether anyone reruns this on a different model family before calling it settled.
Each link below shares sources, entities, or timing with this story.
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
A builder pulled 443 GGUF quantizations across 25 Hugging Face repos and checked whether each file's bits-per-weight matched the type in its name. 64 of them didn't. (r/LocalLLaMA) The mechanism is clean, which is what makes it bad. K-quants and i-quants need the first tensor...
A 179-upvote r/LocalLLaMA thread turned into an impromptu multi-hardware benchmark. OP reports 36 tok/s on a 4GB card where Qwen3.5-9B manages 5 at comparable quality. One commenter ran Q6 CPU-only on an i7-13700K at ~17-20 tok/s. Another ran a sub-4GB Q3 through 17 tool calls...
After the first version drew "Minecraft is in the training data" pushback, the author had the same local Qwen3.8-27B Q4 on a single 4090 add an MLRS system, a rideable skateboard with tricks, an FPV drone, and an in-game computer running a playable game plus an SVG test, coded...
SpeakoFlow Mini fine-tunes Qwen3.5-0.8B to apply only the corrections a speaker actually made and leave the rest alone. On the author's English-only benchmark it scored 70.7% against GPT-5.6 Luna's 65.0% under the same fixed short prompt with reasoning disabled, but the 95% in...
A practitioner running a private trivia set found 3.8 failing questions 3.6 answered reliably, at every quantization and sampling setting tried, then checked Artificial Analysis' Omniscience evaluation and found the same regression in offline no-tool knowledge accuracy. The to...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.