Fetching from the wire…
OSS2026-09-14 · source-backed
anuj0456/OpenArch has single-file implementations for GPT-2 XL, Llama 2/3/4 Maverick, OLMo 2, DeepSeek R1, Gemma 3, Mistral 3, Qwen 3, Kimi K2, GLM 4.5, GPT-OSS and PaliGemma, each making attention type, normalization and positional encoding explicit. Stated goal is clarity over competing with transformers, targeting all 72 architectures in Raschka's gallery. 106 stars, early, and it already covers the models most people deploy.
Each link below shares sources, entities, or timing with this story.
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
The September 11 fine-tuning expansion adds 18 open-weight models (GLM 5.3/5.2/5.1, DeepSeek-V4-Flash variants, Kimi K2.7-Code, Qwen 3.8-27B, Gemma 4) plus Expert LoRA, which puts adapters on the experts themselves. On invented-fact recall, expert-inclusive adapters reached 89...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
DeepClaude hit 470 points on Hacker News. It swaps Claude Code's API backend to DeepSeek V4 Pro while preserving the full agent loop: file editing, bash execution, git tooling, the whole workflow. DeepSeek V4 Pro scores 96.4% on LiveCodeBench at a fraction of Anthropic's prici...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.