Sources
Xiaomi streamed six days of its RL training run live, OOM restarts and cost counter included
Alongside the MiMo-V2.6 weights, Xiaomi ran a public dashboard for six days streaming its reinforcement learning run: 30 steps for each model, $2,620,670 reported for Pro and $854,044 for Flash, roughly 750,000 trajectories, and ~75B and ~81.4B tokens respectively. The log kept its failure notices, including a Pro restart at step 17 from a GPU OOM caused by expert load imbalance, a grader-cluster network failure, and removal of the cyber dataset after bad rollout patterns. The caveat builders should hold onto is that the counters are self-reported and unauditable, and the 30 finished steps are the surviving tail of a longer job whose discarded compute appears in no line item.
Source
↳ Follow the thread