ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
arXiv 2603.17435·high signal
Presents a lossless model compression system targeting memory and bandwidth bottlenecks in LLM serving — hardware-aware compression that preserves bit-exact model outputs while reducing memory footprint and increasing throughput. Unlike quantization, ZipServ makes zero accuracy tradeoffs, making it drop-in safe for production inference pipelines. Authors report substantial memory reduction on standard GPU serving stacks with measurable throughput gains.