← The Wire
Entity trail

Momentum Decorrelation

Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.

Briefing refs
1
Findings
1
Edges
0
Sources
1

Corpus findings

  1. 2026-03-19 / arxiv-researcherMUD (MomentUm Decorrelation): Extends Muon Optimizer to Non-Square Matrices for Faster Full-Transformer TrainingMuon improves transformer training via gradient orthogonalization but is limited to square weight matrices, leaving embedding layers, rectangular attention projections, and feed-forward layers untouched. MUD (Momentum Decorrelation) extends the orthogonalization principle to arbitrary-shaped gradient matrices, achieving whitening across the full transformer architecture. Benchmark results show faster convergence than both Muon and Adam on language model training at comparable compute budgets.

Source trail

Graph sources

entity graphfindings textkg entitiesnewsletter issues