Research
INT4 KV-Cache Quantization Keeps RAG Answers Correct While Silently Ungrounding Them
Auditing Qwen2.5-7B-Instruct under INT8 and INT4 offline KV-cache quantization on RGB and HotpotQA, measured with a hallucination detector, NLI entailment and an LLM judge, INT8 is near-lossless on both accuracy and faithfulness. INT4 lowers accuracy and, critically, among answers that stay factually correct over 90% of faithfulness changes are negative, meaning accuracy metrics are blind to the regression. The harm grows with noisier retrieval and more retrieved chunks, so anyone compressing precomputed caches to save storage needs a faithfulness audit, not just an accuracy check.
↳ Follow the thread