Research
Periodic Row-wise Muon Cuts Optimizer Time 46.9-54.3% on Diffusion Transformers While Keeping the Quality Advantage Over AdamW
Scaling Muon for Diffusion Transformers (arXiv 2608.20818, Aug 21) establishes Muon's scaling behavior on DiTs from 1.3B to 15B parameters, showing a 12.9-19.1% improvement in best generative quality over AdamW that persists across scales. The catch at scale is the 5-step Newton-Schulz iteration run every step plus full-momentum materialization, which offsets the step-efficiency gain. Periodic Row-wise Muon does a full spectral update once every K steps and a cheap row-wise constrained update otherwise, staying within 0.5% of vanilla Muon on 1.3B-4B and improving 4.5% at 9B, while cutting optimizer time 46.9-54.3%, end-to-end step time 15.7-24.3%, and logical communication volume 66.7%.
↳ Follow the thread