Reddit
Qwen Staged Qwen3.8-Flash-Next on ModelScope as a Qwen4 Architecture Preview, 125B Main Params Plus 51B N-Gram Embeddings and 6B Active
The ModelScope card that went up ahead of release described a redesigned multimodal MoE with 125B main-model parameters, an additional 51B of n-gram embeddings, and 6B active per token, stating it is built on the next-generation Qwen4 architecture and released early so the community can prepare for the Qwen4 family. Commenters on the 369-upvote r/LocalLLaMA thread captured screenshots before Qwen edited paragraphs out of the readme minutes later, and noted an FP8 version was listed with no FP4 or QAT-only release. The precedent worth planning around is Qwen3-Next, whose novel architecture took roughly two months to land in llama.cpp.
↳ Follow the thread