Qwen3.8-Flash-Next Got a Release-Day Megathread and Unsloth Day-0 Quants Before the Weights Existed
r/LocalLLaMA moderators pinned a release-day megathread for Qwen3.8-Flash-Next with an estimated drop of 2026-08-26 15:00 UTC, pointing at the Hugging Face and ModelScope pages, while Unsloth's day-0 quant announcement pulled 717 upvotes with a top comment thread of people who had just finished tuning 3.8-27B ('paint is still wet on 27b'). The architecture is the headline: a Qwen4-preview multimodal MoE with GDN gated-delta hybrid layers and Qwen Sparse Attention, reported at 125B main parameters plus 51B N-gram embeddings and 6B active per token, with a claimed ~1/9 the training cost of Qwen3.7-Plus at comparable capability. Commenters are already asking for llama.cpp flags to place sparse KV cache on SSD (n-ssd-ffn, n-cpu-ffn), which is the practical bottleneck for running this locally.
↳ Follow the thread