Four-Bit Quantization Still Reproduces Most Memorized Sequences — and the Surviving Fraction Grows With Model Size
Bits and Memories (arXiv 2607.25451, 2026-07-28) argues that the privacy literature has been measuring the wrong thing: membership inference, rather than verbatim reproduction of training data, which is what practitioners actually worry about. Using Pythia models and their known-memorized sequence sets across five precision levels down to four bits and three model sizes, the authors find quantization is a 'selective forgetter' — verbatim memorization drops faster than perplexity at every precision and size, under two unrelated quantization algorithms and two evaluation corpora. But selectivity is not a defense: at the largest model studied, four-bit quantization still reproduces most memorized sequences while giving up only a few percent of capability, and the surviving memorized fraction increases with model size.
↳ Follow the thread