Dispatch
Qwen3.8-Max-0902 more than doubles TerminalBench 3.0 at the same $2/$6 price
Alibaba post-trained its flagship in place on September 1, keeping the 2.4T-parameter base, 1M context, and $2 input / $6 output per million tokens. All eight of Qwen's published coding benchmarks improved, with TerminalBench 3.0 going 11.3 to 29.0, DeepSWE 1.1 going 56.6 to 69.3, QwenSWEbench V2 going 55.1 to 70.0, and JobBench going 53.4 to 64.0; multimodal scores moved only 0.4 to 3 points. Qwen's own table still puts Claude Opus 5 ahead on most coding sets, and one caveat worth logging is that Alibaba has not mapped the -0902 label to the Qwen3.8-2.4T-A95B open-weight checkpoint, so self-hosters cannot assume they are running the same thing.
↳ Follow the thread