FlashLoop Cuts Looped-Transformer KV Cache Up to 6x Without Retraining
arXiv·low signal
arXiv 2609.29812 finds that the extra compute from looping concentrates on a few tokens and a stable subset of key columns, and that KV residuals between loops quantize well. The training-free framework combines token-sparse updates, sparse attention and KV-residual quantization. The authors report lossless accuracy with up to 1.64x end-to-end speedup and 6x less KV-cache memory.