Fetching from the wire…
Models2026-09-21 · source-backed
Alibaba open-sourced it September 20, unifying text-to-image, editing and native transparent generation in one model: 7B across 32 single-stream DiT layers, paired with a Qwen3-VL 8B text encoder and a 64-channel RGBA VAE at 16x spatial compression. It scores 60.28 on the public open-source leaderboard against Nano Banana 2.0 at 59.82 and GPT Image 1.5 at 59.65, accepts up to 10 reference images, supports bounding-box/brush/mask local edits and native 2K output. The visual generator is down from about 20B in the original Qwen-Image, so frontier-class editing runs on far less VRAM than the closed competitors it beats.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Alibaba post-trained its flagship in place on September 1, keeping the 2.4T-parameter base and 1M context. All eight published coding benchmarks improved: TerminalBench 3.0 from 11.3 to 29.0, DeepSWE 1.1 from 56.6 to 69.3, QwenSWEbench V2 from 55.1 to 70.0, JobBench from 53.4...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
The Max tier is Qwen's flagship proprietary line, distinct from the open-weight Qwen3 series, continuing the Chinese frontier release cadence alongside Moonshot's K3. The practical question for anyone outside China is whether Max-tier access lands on international API endpoint...
Staged on ModelScope for 23:00 Beijing time August 26, it's roughly 125B parameters plus a separate N-gram embedding table of about 51B, activating 6B per token, with GDN gated-delta hybrid layers and Qwen Sparse Attention. Alibaba frames it as a technology preview of the Qwen...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.