Fetching from the wire…
Models2026-09-24 · source-backed
A YOCO-style self-decoder/cross-decoder split paired with token-level sparse attention, where cross-decoder KV caches are built from self-decoder hidden states. On an 80B-A3B MoE it beats HySparse and hybrid SWA on long-context retrieval and multi-turn agent tasks while cutting both prefill compute and KV storage (arXiv). This is the attention design behind MiMo-V3.
Each link below shares sources, entities, or timing with this story.
Xiaomi released Pro and Flash with weights, a technical report, RL environments and training code. Pro scores 46.32 on the Artificial Analysis Intelligence Index and 72.57 on DeepSWE v1.1; Flash gets 65.68. API pricing unchanged at $0.435/$0.87 per million for Pro, $0.14/$0.28...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
This comparison ran the baseline the retrofitted-linear-attention literature skipped. Across multiple LLMs and downstream tasks SWA with sinks matches or beats post-trained linear attention, and on Needle-in-a-Haystack and BABILong it scores 2 to 10 times higher. The recommend...
arXiv 2607.29516, from a Meta-affiliated team including Nachi Nagappan and Peter Rigby, starts from the premise that agents now generate code faster than peer review absorbs it, while existing AI reviewers over-index on style and under-index on correctness, security and perfor...
The stealth model is Xiaomi's MiMo V2 Flash: 309B MoE with 15B active parameters, 256K context, hybrid-thinking toggle. 73.4% SWE-Bench Verified — top open-source globally — approaching GPT-5-High at roughly 3.5% of the cost. A successor model was teased in the same OpenClaw P...
Artificial Analysis ranked it 46 on the Intelligence Index, #1 of 114, ahead of Kimi K3 at 44 and GLM-5.3 at 45, with leading closed models at 53. Natively omnimodal MoE, 1M context, weights on Hugging Face, $0.435 per million input and $0.87 per million output. The companion...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.