VestigeKV Evicts a NoPE KV Cache Using a Vestigial RoPE Branch, Holding 1.00 Retrieval at 32x Compression
Attention-observed selectors like H2O and SnapKV collapse to 0.00-0.33 needle retrieval on a NoPE MLA model because a long-lived cache must be compressed before the queries that will read it exist. On Kimi Linear, VestigeKV evicts by a query-independent signal the cache already carries, the 64-dimensional decoupled branch that NoPE training repurposes from a RoPE vestige into a salience channel, reading 11% of each row and moving non-top rows exactly into a GPU-resident archive rather than deleting them, with no training, quantization or kernel change. Retrieval holds at 1.00 under 8x and 0.92 under 32x from 8k to 65k context, and the effect is NoPE-exclusive: the identical operator on a RoPE MLA collapses to 0.08.
↳ Follow the thread