Research
MoE Routers Share One Geometry Across Depth Once Layer-Specific Coordinates Are Aligned Away
Sparse MoE models parameterize a separate router at every sparse layer, yet routing decisions across depth are partly predictable from earlier signals; this work shows why. Isolating each router's control subspace and aligning the subspaces into a shared canonical frame with generalized orthogonal Procrustes analysis, a single linear transition reaches R²=0.39-0.71 and retains 79-90% of the predictive power of separately fitted layer-specific dynamics. A matched-rank comparison separates this from generic smoothness: residual representations are often easier to predict across layers, but router-control states preserve the model's actual expert choices far more faithfully.
↳ Follow the thread