Fetching from the wire…
Public story · 2026-09-08 · high
It ships under Apache 2.0 with a 131,072-token context and runtime builds for five inference stacks on day one.
Why now: OpenBMB published the model, the score, and every build together on September 7.
OpenBMB released MiniCPM5-2B on September 7. The 2-billion-parameter model posted the top Intelligence Index score for any open-weight model under 4B parameters.
That score, 15 on Intelligence Index v4.2, matters for anyone picking a small model to run on a laptop or a cheap GPU instance. OpenBMB's model card puts the average at 53.9, calling it open-source state of the art for the 2B class.
The architecture is a 42-layer dense model with grouped-query attention, running in BF16, with a 131,072-token context window. Post-training used 400 billion tokens of supervised fine-tuning plus reinforcement learning, with specialized teachers and on-policy distillation pulled from 16 expert models. That's a heavier training recipe than most 2B releases bother with.
GGUF, MLX, GPTQ, vLLM and SGLang builds arrived alongside the release, plus the UltraData training sets. The model is Apache 2.0 licensed, so nothing in the card or the weights blocks commercial use. Five runtime formats on day one means it's already installable on a Mac, a consumer GPU, or a production inference server. There's no wait for the community to backport quantizations.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
Lily is Apache 2.0 inside pplx-garden, a small Metal inference server built for Qwen3.6-35B-A3B converted to MLX affine 4-bit, exposing a minimal OpenAI-compatible chat API with greedy decoding. The README explicitly rules out dense and smaller Qwen checkpoints, BF16, GGUF, AW...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.