Fetching from the wire…
Models2026-09-15 · source-backed
Someone did the arithmetic the leaderboard hides, at Q4_K_M with no drafter, no vision, 128K KV cache (r/LocalLLaMA). K2 Horizon 36B-A4B needs 2 GiB dense plus 19 GiB experts plus 6.7 GiB context. The 7B needs 5.2 GiB weights plus 5 GiB context. Compare Qwen3.6-35B-A3B at 0.7 GiB of context. The conclusion: 36B-A4B only makes sense on exactly 16GB VRAM with 32GB host RAM, and on 24GB VRAM Qwen3.8-27B is faster, smarter and fits 256K context. Parameter count is the wrong x-axis for anyone choosing a local model.
Each link below shares sources, entities, or timing with this story.
An r/LocalLLaMA thread asking where the promised MoE went turned up a hard artifact: modelscope/ms-swift commit a45f1d4, titled "fix wrong model-ids," removes the Qwen/Qwen3.8-35B-A3B and -FP8 entries from swift/model/models/qwen.py and substitutes the dense 27B. Why it matter...
IFM's lineup showed up on the Artificial Analysis Intelligence Index with the 7B slotting between two much larger Qwen models (r/LocalLLaMA). Two corrections from the thread and the model card: the HF config reports about 8.999B parameters in BF16 for a dense model, and the RE...
An r/LocalLLaMA post (215 upvotes) ran identical inputs through Qwen 35B-A3B and Gemma 26B-A4B and found the tokenizers diverge almost entirely on code: 2.6x apart on HTML/JS, but 1,025 vs 1,039 tokens on a 55-line instruction document. Near-identical on prose. That's a near-3...
An r/LocalLLaMA post reports a working deployment on 80x RTX 5090 connected over 25 gigabit Ethernet rather than NVLink or InfiniBand — roughly 2.5TB aggregate VRAM against a ~594GB MXFP4 weight file, surplus absorbed by activation and KV overhead. Expert-parallel MoE over com...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
A practitioner running a private trivia set found 3.8 failing questions 3.6 answered reliably, at every quantization and sampling setting tried, then checked Artificial Analysis' Omniscience evaluation and found the same regression in offline no-tool knowledge accuracy. The to...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.