Google TurboQuant: 6x KV-Cache Compression With Zero Accuracy Loss — ICLR 2026 Paper
r/LocalLLaMA / Google Research·medium signal
Google Research published TurboQuant, a compression algorithm that reduces LLM key-value cache memory by at least 6x and delivers up to 8x inference speedup with zero accuracy loss. The technique combines Quantized Johnson-Lindenstrauss transforms with PolarQuant for high-quality compression. Published as an ICLR 2026 conference paper, this directly impacts local model deployment economics and semantic search at scale.