Research
Trainable Rotary Positional Embeddings: Making RoPE's Rotation Manifold a Learnable Expressivity Dimension
Cheng, Sun, and Lu argue that RoPE's rotation manifold has been treated as fixed hand-crafted structure populated only by discrete ordinal indices, while semantic embedding space gets enormous learning capacity. They propose making the rotation space learnable with temporal and semantic rotary encodings, positioning this as a 'largely overlooked second dimension of expressivity' in attention mechanisms. If validated at scale, this could systematically improve any Transformer using RoPE.
Source
↳ Follow the thread