Practitioner Report: DeepSeek V4-Flash-0731 Falls Apart Under Quantization in a Way the Preview Build Did Not
An r/LocalLLaMA weekend writeup (189 upvotes, 86 comments) reports that quantization 'hits this thing like a truck' — Q2 and Q3 weights of V4-Flash-0731 behave like a different model, reasoning traces change shape, and results land a full tier below the officially served version. The practical thresholds given: Q3 finally replaces Qwen3.6-27B and pulls ahead in large repos and Claude Code-style harnesses with 30k+ token system prompts, while Q2 loses outright to Qwen3.6-27B at Q8. A commenter posted Unsloth KL-divergence charts showing the earlier DS4 Flash preview was forgiving of quantization while 0731 is not, with poor KLD even at IQ4_XS — and the poster flags the model is weak on general knowledge, fine if you're tool-calling but not for airgapped use.
↳ Follow the thread