Fetching from the wire…
Infra2026-09-13 · source-backed
ilintar published the full optimization writeup after a closed-source server called Halogen claimed 1.2k t/s prefill on Qwen3.8 Flash Next while the community fork sat near 400. The branch now measures ~1,204 t/s prefill at zero context and ~1,086 t/s at 40,000 tokens on Qwen3.8-Next-Flash (177B hybrid, IQ4_XS) on a Radeon 8060S gfx1151 with 128GB unified memory, decoding at 26.28 ± 0.29 t/s. The wins are retained PM4 command lists for HIP graphs so packets aren't re-encoded per launch, plus ROCm-side flash attention, MMA and MoE routing work. The author expects the approach to carry to GLM 5.3 Flash's sparse attention. 190 upvotes on r/LocalLLaMA.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Claims up to 2x faster generation, and adds fine-tuning of both MoE models on text or image datasets on Apple Silicon via MLX (release). Follow-up turns in long Qwen chats on Mac are reported up to 30x faster, MLX models now use full context size, and GLM-5.3 MLX fine-tunes ex...
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen...
Created August 24, it holds a trendingScore of 3,967 against second-place GLM-5.3-Flash at 1,376 (Hugging Face). The near-1:1 like-to-download ratio means almost everyone bookmarking it hasn't pulled weights, and the unsloth GGUF conversion at 4,354 downloads is absorbing comp...
Alibaba published a fine-grained MoE with 2.4T total / 95B active, 512 experts, and a 92-layer hybrid full/linear attention backbone. vLLM shipped day-0 support verified on NVIDIA and AMD with ready 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for...
A month of solo work produced quants for LongCat-Flash-Lite-Sparse, Qwen3.8-27B, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with vision. Getting there meant writing Heretic support for the architecture from scratch and then adding llama.cpp support, and main...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.