Tools
Qwen releases Qwen3.8-Flash-Next weights as an early architecture preview of Qwen4, and ollama shipped MLX support two days later
QwenLM/Qwen3.8-Flash-Next appeared on GitHub 2026-08-24 and is at 179 stars, opening the weights of a multimodal MoE that Qwen explicitly frames as an early preview of the Qwen4 architecture, the same role Qwen3-Next played for Qwen3.5. The hybrid Gated DeltaNet plus Gated Attention design it previews has already carried through the Qwen3.5 to Qwen3.8 series. The release is corroborated downstream: ollama v0.33.1 (2026-08-26) lists MLX Qwen3.8 Flash Next support, and transformers v5.16.0 added the related Qwen4-Exp with Qwen Sparse Attention and Per-Layer Embedding.
Source
↳ Follow the thread