Fetching from the wire…
Public story · 2026-07-25 · high
The model card includes full checkpoints and training code, but the license is stricter than the Apache terms on AMD's earlier 3B release.
Why now: This is AMD's most complex Instella release yet, the first at mixture-of-experts scale, making it the first real test of whether Instinct chips can carry a training run this large alone.
AMD uploaded Instella-MoE-16B-A3B-Think to Hugging Face, a 16 billion parameter mixture-of-experts model trained entirely on its own MI300X and MI325X chips. That's the company's clearest proof yet that a frontier-scale mixture-of-experts model can be trained start to finish without a single Nvidia GPU.
The model runs 64 experts, routing 6 per token plus 2 shared, with 2.8 billion parameters active per token. It uses Gated Multi-head Latent Attention and FarSkip-Collective connectivity, all built on AMD's own Primus training framework.
AMD published the full training recipe alongside the weights: data mixtures, intermediate checkpoints from pretraining through reinforcement learning, and inference code.
The catch, flagged by r/LocalLLaMA: the license is ResearchRAIL, academic and research use only. AMD's earlier Instella-3B shipped under Apache-style terms open to commercial use. This one you can study, not ship.
My read: the strict license is the tell. AMD built this to show MI300X buyers and researchers that the chips can carry a training run this complex, not to hand developers a model they can ship. Whether AMD's next release comes with looser terms is the thing worth watching.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
Alibaba published a fine-grained MoE with 2.4T total / 95B active, 512 experts, and a 92-layer hybrid full/linear attention backbone. vLLM shipped day-0 support verified on NVIDIA and AMD with ready 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
3B active parameters, beats Qwen3.5-35B-A3B on AIME 2025 (92.4 vs 91.9), LiveCodeBench v6 (87.2 vs 74.6), and surpasses the larger Nemotron-3-Super-120B. Available on Ollama and HuggingFace under open license. Source
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.