Research
KV Cache Tiering Buys 73x More Sessions Per GPU, and the Eviction Policy Barely Matters
Where Should the KV Cache Live? (arXiv 2609.16215, submitted 14 Sep 2026) simulates GPU HBM, CPU DRAM, and SSD tiers calibrated against a random forest execution-time predictor, comparing recency, reuse frequency, predicted reuse, and an EWMA predictor with prefetch lookahead across chat, agent, and document QA workloads. Tiering supported 73.02x more concurrent sessions per GPU and cut cost per session 62.04x, but the authors attribute the gains to tier capacities of 1 + 8 + 64, not to placement policy. Because decode was compute bound at batch size one in their setup, policy mainly changed PCIe migration traffic and time to first token, with recency producing 2.30x less migration traffic.
↳ Follow the thread