Relation Replaces Attention as the Token-Mixing Primitive and Beats MHA on Validation NLL at 10M, 30M and 100M
Relation reorders the attention computation, first organizing pairwise evidence into explicit Self and Exchange relations and only then deriving information flow, which yields a family of variants including FlashRelation, Linear Relation, Hybrid Relation and a KV-style Relation Cache. Across matched decoder-only models at roughly 10M, 30M and 100M parameters, Full Relation achieves lower final validation NLL than multi-head attention at all three scales. FlashRelation runs 3.60-4.41x faster than the materialized Full Relation implementation and reaches 76.4-84.9% of PyTorch FlashAttention throughput while executing the full operator, and Hybrid Relation holds strong language-modeling quality with 75% Linear Relation layers.
↳ Follow the thread