AMD Quietly Uploads Instella-MoE-16B-A3B-Think — 64 Experts, MI300X-Trained, but Under a Research-Only License
AMD has posted Instella-MoE-16B-A3B-Think to Hugging Face: 16B total parameters with 2.8B active per token, 64 experts routing 6 per token plus 2 shared experts, using Gated Multi-head Latent Attention and 'FarSkip-Collective' connectivity. It was trained entirely on AMD Instinct MI300X and MI325X hardware via AMD's Primus framework, and AMD is releasing the full training recipe — frameworks, data mixtures, intermediate checkpoints from pretraining through RL, and inference code. The catch that r/LocalLLaMA flagged (153 upvotes) is the license: ResearchRAIL, restricted to academic and research use, a real step back from the Apache-style terms on AMD's earlier Instella-3B line. Strategically this is AMD demonstrating it can train a competitive MoE end-to-end without Nvidia silicon; commercially, the license means you cannot ship it.
↳ Follow the thread