Music Transformers Get Less Equivariant as They Scale, Spending Capacity Memorizing Absolute Patterns
Humans recognize a passage shifted in time or transposed in pitch, but the authors' analysis shows standard music transformers map such inputs onto uncorrelated representations — and get progressively worse at it with larger size and longer training, implying extra capacity goes to memorizing absolute patterns rather than shared structure. The Equivariant Music Transformer enforces equivariance via self-distillation, jointly optimizing next-token prediction with an auxiliary equivariance regularization loss. The auxiliary loss acts as a beneficial regularizer that improves next-token prediction at the same time, beating data augmentation, feature engineering, and SOTA baselines on both objective and subjective evaluation.
↳ Follow the thread