Dispatch
Qwen3.8-Flash-Next is a 125B/6B-active multimodal MoE and an early look at the Qwen4 architecture
Qwen released the open-weights model on 2026-08-26, describing it as a multimodal mixture-of-experts with 125 billion total parameters and 6 billion active, explicitly framed as a preview of the architecture Qwen4 will use. Simon Willison ran quantized Unsloth builds on a DGX Spark, a 72.5GB UD-IQ1_S and a 78.9GB UD-Q2_K_XL, both on Hugging Face, and got his best result from UD-Q2_K_XL at xhigh reasoning effort. The 6B active count is the number that matters for local inference: it puts a 125B multimodal model inside a single workstation's reach.
↳ Follow the thread