Fetching from the wire…
Models2026-09-24 · source-backed
Five models covering ASR, TTS and Realtime. ASR-Next adds speaker, timestamp, emotion and background-sound analysis; TTS-Next generates voices, effects and ambient audio together. Roughly 70% off TTS, 85% off Realtime, up to 95% off ASR. Standard ASR covers 30 languages plus 16 Chinese dialects at about 160ms to first character (The Decoder). At that price, high-volume voice agent economics change shape.
Each link below shares sources, entities, or timing with this story.
Alibaba post-trained its flagship in place on September 1, keeping the 2.4T-parameter base and 1M context. All eight published coding benchmarks improved: TerminalBench 3.0 from 11.3 to 29.0, DeepSWE 1.1 from 56.6 to 69.3, QwenSWEbench V2 from 55.1 to 70.0, JobBench from 53.4...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
The Max tier is Qwen's flagship proprietary line, distinct from the open-weight Qwen3 series, continuing the Chinese frontier release cadence alongside Moonshot's K3. The practical question for anyone outside China is whether Max-tier access lands on international API endpoint...
25,000 fake accounts. 28.8 million Claude conversations. Six weeks. And the thing they were harvesting wasn't trivia, it was software engineering and agentic reasoning. In a June 24 letter to US senators and the White House, Anthropic alleged that operators tied to Alibaba's Q...
The winner isn't the story. The methodology is. Databricks published its internal coding-agent benchmark: real engineering tasks pulled from its own multi-million-line codebase spanning Python, Go, TypeScript, and Scala. Roughly 25% low-complexity tasks, about 60% medium. Not...
Staged on ModelScope for 23:00 Beijing time August 26, it's roughly 125B parameters plus a separate N-gram embedding table of about 51B, activating 6B per token, with GDN gated-delta hybrid layers and Qwen Sparse Attention. Alibaba frames it as a technology preview of the Qwen...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.