Reddit
Alibaba Releases Qwen3.8-2.4T-A95B Open Weights — 2.4 Trillion Parameters, 512 Experts, With Day-0 vLLM and SGLang Support
Alibaba published open weights for Qwen3.8-2.4T-A95B, a fine-grained MoE with 2.4T total parameters, 95B active per token, 512 experts and a 92-layer hybrid full/linear attention backbone — one of the largest open-weight models ever released. vLLM shipped day-0 support verified on both NVIDIA and AMD with ready-made 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for 8xMI355X), and NVIDIA measured over 4K tokens/sec per GPU on GB300 NVL72. The top r/LocalLLaMA post of the day at 1,507 upvotes, though the HuggingFace discussion tab is dominated by complaints that the open release is text-only and stripped of the 1M context and vision that Qwen3.8-Max has behind the API.
↳ Follow the thread