Fetching from the wire…
Models2026-08-02 · source-backed
Released July 31 under South Korea's Sovereign AI Foundation Model Project: hybrid-attention MoE, 750B total with ~37B active per token, 256 experts with 8 selected, a 262,144-token context window, 10-language coverage. Over 3x the size of the first K-EXAONE (236B/23B), lifting the average benchmark score from 63.3 to 70.1, with 83.5 MMLU-Pro and 92.3 AIME 2026. Apache 2.0 makes it one of the largest permissively-licensed open-weight models available. Sovereign AI programs producing genuinely competitive open weights is a 2026 development I did not expect at the start of the year.
Each link below shares sources, entities, or timing with this story.
GitHub | 196B total / 11B active, Apache 2.0 StepFun's sparse MoE activates only 11B of 196B parameters per token, delivering 74.4% SWE-bench, 97.3% AIME 2025, and 100-300 tok/s throughput. Supports INT4 GGUF for local inference. Apache 2.0 licensed. One of the most capable fu...
Alibaba's 9B model outperforms OpenAI's 120B on GPQA Diamond (81.7 vs 71.5), MMLU-Pro (82.5 vs 80.8), and multilingual MMMLU. Uses hybrid Gated Delta Network + sparse MoE with 262K native context. Apache 2.0 on HuggingFace. VentureBeat
MAI-Code-1-Flash, a 5B-parameter coding model, is in GitHub Copilot and VS Code, and Microsoft says it beats Claude Haiku 4.5 across core coding benchmarks, +16 points on SWE-Bench Pro at 51.2% versus 35.2%, using up to 60% fewer tokens. MAI-Thinking-1, a 35B-active MoE with a...
Moonshot AI dropped Kimi K2.7-Code on Hugging Face on June 12. The specs are loud: 1T-parameter MoE with 32B active across 384 experts, a 256K context window, Modified MIT license, tuned for long-horizon agentic software engineering (MarkTechPost). Moonshot reports +21.8% on K...
Google released Gemma 4 on April 2 with four model variants: E2B, E4B, 26B MoE, and 31B Dense. The license change is the first thing worth noting. Every previous Gemma had restrictions that made lawyers nervous. Gemma 4 is Apache 2.0. Full stop. Use it in any product, any way...
Can a model that fits on a Raspberry Pi do reliable tool calling? Two independent labs just answered yes. PrismML emerged from stealth March 31 with Bonsai, the first commercially viable 1-bit LLMs built on Caltech research. The 8B model fits in 1.15GB (vs 16GB for FP16), runs...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.