Fetching from the wire…
Public story · 2026-07-25 · high
The model card includes full checkpoints and training code, but the license is stricter than the Apache terms on AMD's earlier 3B release.
Why now: This is AMD's most complex Instella release yet, the first at mixture-of-experts scale, making it the first real test of whether Instinct chips can carry a training run this large alone.
AMD uploaded Instella-MoE-16B-A3B-Think to Hugging Face, a 16 billion parameter mixture-of-experts model trained entirely on its own MI300X and MI325X chips. That's the company's clearest proof yet that a frontier-scale mixture-of-experts model can be trained start to finish without a single Nvidia GPU.
The model runs 64 experts, routing 6 per token plus 2 shared, with 2.8 billion parameters active per token. It uses Gated Multi-head Latent Attention and FarSkip-Collective connectivity, all built on AMD's own Primus training framework.
AMD published the full training recipe alongside the weights: data mixtures, intermediate checkpoints from pretraining through reinforcement learning, and inference code.
The catch, flagged by r/LocalLLaMA: the license is ResearchRAIL, academic and research use only. AMD's earlier Instella-3B shipped under Apache-style terms open to commercial use. This one you can study, not ship.
My read: the strict license is the tell. AMD built this to show MI300X buyers and researchers that the chips can carry a training run this complex, not to hand developers a model they can ship. Whether AMD's next release comes with looser terms is the thing worth watching.
Each link below shares sources, entities, or timing with this story.
AMD partners with OpenAI / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (AMD partners with OpenAI); both cover Apache, LocalLLaMA, MoE, NVIDIA; reported by the same outlet (huggingface.co).
NVIDIA released Nemotron / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (NVIDIA released Nemotron); both cover A3B, LocalLLaMA, NVIDIA; overlapping topics (active, license).
Anthropic uses NVIDIA / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Anthropic uses NVIDIA); both cover LocalLLaMA, MoE; reported by the same outlet (huggingface.co).
AMD partners with Meta / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (AMD partners with Meta); both cover Apache, LocalLLaMA, MoE; reported by the same outlet (huggingface.co).
AMD partners with Meta / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (AMD partners with Meta); both cover Apache, MoE; reported by the same outlet (huggingface.co).
NVIDIA released Blackwell / Shared entities / Earlier coverage / Downstream implication
Linked by a graph relationship (NVIDIA released Blackwell); both cover AMD, LocalLLaMA, NVIDIA; earlier AMD coverage from 2026-04-10.
NVIDIA benchmarked against DeepSeek-R1 / Shared entities / Earlier coverage
Linked by a graph relationship (NVIDIA benchmarked against DeepSeek-R1); both cover Apache, LocalLLaMA, MoE; earlier Apache coverage from 2026-04-05.
Hermes Agent supports NVIDIA / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Hermes Agent supports NVIDIA); both cover LocalLLaMA, Think; reported by the same outlet (huggingface.co).