Hacker News
Xiaomi Puts a Live RL Post-Training Dashboard for MiMo-V2.6 on the Public Internet
MiMo team head Luo Fuli posted a link on X on September 17 to a minute-by-minute dashboard of MiMo-V2.6's reinforcement learning run, showing accepted and judged sample counts, pass rate (around 0.66 on n≈8,100-8,500), remaining and partial samples, prewarm and step numbers, plus token consumption and training cost. The run started 2026-09-15 10:32 UTC and Luo says the team spent six months on how far RL scales, expanding compute, training environments and task systems, and reward evaluation simultaneously. No other lab publishes a live post-training telemetry feed, which makes this a free reference for anyone calibrating their own RL run economics.
↳ Follow the thread