Fetching from the wire…
Public story · 2026-08-25 · high
The pulled card listed 125B main-model parameters, 51B of n-gram embeddings, and 6B active per token before Qwen removed the details.
Why now: Qwen pulled the architecture details from the listing within minutes, so the screenshots commenters saved before the edit are the only public record of the original numbers.
Qwen listed Qwen3.8-Flash-Next on ModelScope with a model card describing it as an early preview of the company's next-generation Qwen4 architecture. Qwen edited out that description within minutes. Commenters had already saved screenshots of the original wording.
The card put real numbers on a Qwen4-class design for the first time. It described a redesigned multimodal mixture-of-experts model: 125B main-model parameters, an additional 51B in n-gram embeddings, and 6B active per token. Qwen said it was releasing the preview early so developers could start preparing for the architecture change.
The listing included an FP8 version. There was no FP4 build and no QAT-only release, so anyone planning to run this at lower precision doesn't have an official option yet.
The precedent here is Qwen3-Next, the last time Qwen introduced a new model architecture instead of an update to an existing one. That model took about two months to reach llama.cpp support after its release. The edited Qwen3.8-Flash-Next card didn't include a release date, just the preview.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Alibaba released Qwen); both cover ModelScope, Qwen, Qwen3; overlapping topics (architecture, around, card).
Linked by a graph relationship (Alibaba released Qwen); both cover ModelScope, MoE, Qwen, Qwen3; overlapping topics (active, around, qwen).
Ollama supports Qwen / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Ollama supports Qwen); both cover MoE, Qwen, Qwen3; earlier MoE coverage from 2026-04-23.
Alibaba released Qwen / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Alibaba released Qwen); both cover MoE, Qwen, Qwen3; earlier MoE coverage from 2026-04-21.
Qwen competes with Meta / Shared entities / Earlier coverage
Linked by a graph relationship (Qwen competes with Meta); both cover Flash, MoE, Qwen; earlier Flash coverage from 2026-08-16.
Linked by a graph relationship (Qwen competes with Meta); both cover Flash, MoE, Qwen3; earlier Flash coverage from 2026-08-10.
Linked by a graph relationship (Qwen competes with Meta); both cover MoE, Qwen, Qwen3; earlier MoE coverage from 2026-04-02.
Ollama supports Qwen / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Ollama supports Qwen); both cover MoE, Qwen; overlapping topics (active, community, qwen).