First Systematic Map of Massive Activations in Hybrid Linear-Attention LLMs: Pre-Attention Spikes and Inter-Spike Plateaus, 1.2B to 397B Parameters
arXiv 2608.12149 (August 12, 2026) reports that in layer-interleaved hybrid linear attention models, massive activations spike immediately before full-attention layers (pre-attention spikes) and can persist through intervening linear-attention layers as inter-spike plateaus — and as full attention gets denser, the spikes connect into the stable morphology familiar from full-attention LLMs. The pattern recurs across five linear-attention architectures, six hybridization configurations, five data domains and open models from 1.2B to 397B total parameters. Controlled GDN-hybrid pretraining up to 1.3B shows full-attention output gating strongly attenuates magnitudes without changing layerwise organization, which matters for anyone quantizing these hybrids.
↳ Follow the thread