Fetching from the wire…
Models2026-09-22 · source-backed
The same Apsara announcement describes a month of fully automated runs covering pipeline design, data validation, iterative experimentation and error diagnosis with no human in the loop, plus a separate 60-hour chip-design run producing production-grade bus modules with a 42% area reduction. Vendor self-report, no third-party replication, so treat it as a claim. It's still the most concrete recursive-self-improvement number any major lab has attached a figure to. (Pan African Visions)
Each link below shares sources, entities, or timing with this story.
Alibaba's model was reported best overall on Artificial Analysis' Agentic Index, drawing 540 points on HN. Readers watching the page saw Qwen at 55.4 vs Opus Max at 55.3, then on reload Opus Max at 59.2 vs Qwen at 58.4. George from Artificial Analysis replied in-thread that th...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
The letter to Senators Tim Scott and Elizabeth Warren, dated June 10 and surfacing publicly this week, frames it as model distillation run against Claude at scale (Anthropic). A related claim pegs it at 28.8 million fraudulent exchanges, though that figure is single-sourced an...
Alibaba post-trained its flagship in place on September 1, keeping the 2.4T-parameter base and 1M context. All eight published coding benchmarks improved: TerminalBench 3.0 from 11.3 to 29.0, DeepSWE 1.1 from 56.6 to 69.3, QwenSWEbench V2 from 55.1 to 70.0, JobBench from 53.4...
Alibaba published a fine-grained MoE with 2.4T total / 95B active, 512 experts, and a 92-layer hybrid full/linear attention backbone. vLLM shipped day-0 support verified on NVIDIA and AMD with ready 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.