Fetching from the wire…
Public story · 2026-08-16 · high
Tested across models from 1.2 billion to 397 billion parameters, the paper says to check this before quantizing a hybrid.
Why now: This reached the August 16 coverage as hybrid linear-attention models are already shipping as open weights across the exact range it tested.
Massive activations spike right before full-attention layers in hybrid linear-attention models, then hold steady through the layers between, per a paper on arXiv.
That's a caution for anyone compressing or quantizing a hybrid linear-attention model, per the paper's own advice: check where these spikes sit before you compress. The pattern held from 1.2 billion parameters up to 397 billion, so this isn't a one-model quirk.
As the ratio of full-attention layers gets denser, the spikes connect into the same stable shape seen in ordinary full-attention language models. The paper calls the calm stretches between spikes inter-spike plateaus. It tested five linear-attention architectures, six hybridization configurations, and five data domains to confirm the pattern holds.
Controlled pretraining of GDN-hybrid models up to 1.3 billion parameters tested whether the spikes can be tuned down. Full-attention output gating cut their size sharply but didn't move where they sit in the network.
Compress a hybrid linear-attention model without tracking which layers sit next to full attention, and you compress blind on the layers this paper flags. That's the paper's own bottom line for anyone about to quantize one.
This reached the August 16 coverage as hybrid linear-attention models are already shipping as open weights across the exact range it tested.
Each link below shares sources, entities, or timing with this story.
Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage / Tension
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (massive, model).
Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (activation, architectur).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (data, model).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (data, layer).
Both cover LLMs; reported by the same outlet (arxiv.org); overlapping topics (layer, model).
Shared entity: LLMs / Same source domain / Earlier coverage / Tension
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-08-07.
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-07-31.
Both cover LLMs; reported by the same outlet (arxiv.org); earlier LLMs coverage from 2026-07-28.