Fetching from the wire…
Public story · 2026-09-02 · high
The September 1 refresh keeps the same price and parameters, but Code Arena now ranks it above Claude on one web-dev test.
Why now: Alibaba posted the update and the benchmark tables on September 1 and 2, 2026.
Alibaba re-trained its flagship Qwen3.8-Max model in place on September 1, keeping the same 2.4 trillion parameter base and $2/$6 price. The update more than doubles the TerminalBench 3.0 score, from 11.3 to 29.0, the largest jump among eight published coding benchmarks.
DeepSWE 1.1 rose from 56.6 to 69.3. QwenSWEbench V2 climbed from 55.1 to 70.0, and JobBench went from 53.4 to 64.0, per CellCog's breakdown. Multimodal scores barely moved, up 0.4 to 3 points depending on the test.
The number getting the most attention is Code Arena's WebDev leaderboard, where the model's score rose from 1,669 to 1,691. That's nominally first place, and the first time a Chinese model has out-scored Claude on that board, according to TechNode's report.
The top comment on r/ClaudeAI pushes back on that number. Code Arena's WebDev test leans heavily on one-shot UI and SVG generation, a narrow slice of what coding involves. Alibaba's own published comparison table backs that up. Claude Opus 5 still leads on most of the coding benchmarks in the set.
There's a second gap for anyone trying to run this locally. Alibaba hasn't said whether the -0902 snapshot maps to an open-weight checkpoint. Self-hosters have no way to confirm they'd be running the same model that produced these scores. For a lab competing partly on openness, that mapping is worth watching for in the next release notes.
Each link below shares sources, entities, or timing with this story.
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
25,000 fake accounts. 28.8 million Claude conversations. Six weeks. And the thing they were harvesting wasn't trivia, it was software engineering and agentic reasoning. In a June 24 letter to US senators and the White House, Anthropic alleged that operators tied to Alibaba's Q...
The Max tier is Qwen's flagship proprietary line, distinct from the open-weight Qwen3 series, continuing the Chinese frontier release cadence alongside Moonshot's K3. The practical question for anyone outside China is whether Max-tier access lands on international API endpoint...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.