Reddit
r/LocalLLaMA Is Reporting That q8 KV Cache Measurably Hurts Qwen3.8-27B, Not Just Compresses It
A post with 53 upvotes but 74 comments, the highest comment-to-score ratio on the sub today, argues that q8 KV cache quantization degrades Qwen3.8-27B's output quality rather than being the near-free win it is usually treated as. The high engagement ratio matters more than the score here: it is the shape that has previously marked genuinely contested technical claims on this sub. For anyone running 27B locally under VRAM pressure, this is the parameter worth re-testing before trusting the default advice to always enable it.
Source
↳ Follow the thread