Fetching from the wire…
Models2026-09-16 · source-backed
ByteShape released full ShapeLearn GGUFs claiming 3.84 bpw reaches 99.63% of BF16's aggregate score across eight benchmarks and 3.23 bpw reaches 98.72%, with all five models on the measured quality/speed-bpw frontier across six GPUs. The load-bearing argument is methodological: Unsloth Dynamic V3's UD-IQ3_S posted about 20% lower KLD at comparable size yet lost on downstream task results. If you pick quants by KLD alone you may be picking the wrong one.
Each link below shares sources, entities, or timing with this story.
TAK builds an imatrix from a task-specific corpus, finds the smallest size before collapse, then promotes and demotes tensors within a byte budget. No pruning, no fine-tuning, no merging. Held-out reasoning: 82.81% against 83.59% for BF16 and 77.34% for byte-matched Unsloth UD...
An RTX 5080 owner ranked community quantizations by mean KLD and same-top-p agreement rather than a public benchmark. bartowski/Qwen3.8-27B-IQ4_XS won overall, huihui-ai's abliterated UD-IQ4_XS was the best uncensored option, and jpetrina's IQ4_XS-pure is the pick when you nee...
Quesma ran the model across GPQA Diamond, IFBench and Terminal-Bench 2.1 (89 agentic coding tasks) on L40S, H100 and H200 via Modal. Q4_K_M at 17 GB matched BF16 at 55 GB within a point on all three. UD-Q2_K_XL at 10.7 GB held instruction-following but dropped Terminal-Bench f...
Mini is a 35B/~3B-active multimodal MoE with 262K context, 256 routed experts (8 active plus a shared expert) and ~70GB of BF16 safetensors, deployable on two H100s. Pro at 397B scores 56.4 on OSWorld-2 against Qwen3.8-Max at 46.7, plus OSWorld-G 87.4, OSWorld-Verified 82.2 an...
The Quantization-Aware Healing post claims a compressed 4-bit model outperforms the uncompressed one. Technical readers clarified the actual claim: a 120B cut to a 60B BF16 model, then to 60B mxfp4 that beats the 60B BF16 but not the 120B base. Different claim entirely. The sh...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.