Reddit
Xiaomi also shipped MiMo-V2.6-Distill-Qwen-9B so small labs can reproduce the RL work
Alongside Pro and Flash, Xiaomi released MiMo-V2.6-Distill-Qwen-9B, a distillation of the frontier model into a Qwen 9B body, plus a research package that datastudios.org reports contains more than 7,000 reinforcement-learning tasks covering vulnerability reproduction and knowledge work. The 9B distill is explicitly framed for smaller-scale RL research rather than production. Two separate r/LocalLLaMA threads on it drew 251 and 143 upvotes, with commenters treating the RL task set, not the flagship weights, as the more reusable artifact.
↳ Follow the thread