Fetching from the wire…
Models2026-09-05 · source-backed
An RTX 5080 owner ranked community quantizations by mean KLD and same-top-p agreement rather than a public benchmark. bartowski/Qwen3.8-27B-IQ4_XS won overall, huihui-ai's abliterated UD-IQ4_XS was the best uncensored option, and jpetrina's IQ4_XS-pure is the pick when you need context headroom. The bottom of the table is the useful part: a QAT q2_0 at 8.2GiB carries 0.89 mean KLD and only 85.7% top-p agreement, so the cheapest quants are much worse than their file size implies (r/LocalLLaMA).
Each link below shares sources, entities, or timing with this story.
SpeakoFlow Mini fine-tunes Qwen3.5-0.8B to apply only the corrections a speaker actually made and leave the rest alone. On the author's English-only benchmark it scored 70.7% against GPT-5.6 Luna's 65.0% under the same fixed short prompt with reasoning disabled, but the 95% in...
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
The video post shows the 80GB model generating on a 12GB mid-range handset with aggressive quantization on the dense part plus unnamed optimizations. The CPU hits 80C, which the top commenter flags immediately, so this demonstrates what fits rather than something you'd leave r...
Unsloth's UD-Q2_K_XL (78.9 GB) plus a 358,400-token slot via YaRN from the native 262,144 with fp16 KV fits under the default 96GB Metal wired limit, no sysctl hack. Cold prefill runs 1,561 t/s at 5.6K context down to 318 t/s at 111K; a normal incremental turn is 77-854 t/s ou...
189 upvotes on r/LocalLLaMA, with the blunt summary that quantization "hits this thing like a truck." Q2 and Q3 behave like a different model with differently-shaped reasoning traces. Thresholds given: Q3 finally beats Qwen3.6-27B in large repos and harnesses with 30k+ token s...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.