Fetching from the wire…
Public story · 2026-03-19 · source-backed
Muon's gradient orthogonalization improves training but is limited to square weight matrices. MUD (Momentum Decorrelation) extends this to arbitrary-shaped gradient matrices, achieving whitening across embeddings, rectangular attention projections, and feed-forward layers. Faster convergence than both Muon and Adam at comparable compute budgets. arXiv
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / What happened next
Both cover Adam, Muon; reported by the same outlet (arxiv.org); picks up the Adam thread on 2026-08-06.
Shared entity: Muon / Same source domain / What happened next / Downstream implication
Both cover Muon; reported by the same outlet (arxiv.org); picks up the Muon thread on 2026-06-23.
Shared entity: Faster / Shared topic / What happened next
Both cover Faster; overlapping topics (compute, faster); picks up the Faster thread on 2026-04-27.
Adam uses Anthropic / Same source domain
Linked by a graph relationship (Adam uses Anthropic); reported by the same outlet (arxiv.org).
Linked by a graph relationship (Adam uses Anthropic); reported by the same outlet (arxiv.org).
Linked by a graph relationship (Adam uses Anthropic); reported by the same outlet (arxiv.org).
Linked by a graph relationship (Adam uses Anthropic); reported by the same outlet (arxiv.org).
Linked by a graph relationship (Adam uses Anthropic); reported by the same outlet (arxiv.org).