Reddit
Kaitchup's Qwen3.8 27B quant sweep puts UD Q3_K_XL at 100% accuracy in 12.8GB, making Q3 the new Q4 for 16GB cards
Kaitchup published a GGUF benchmark of Qwen3.8 27B across quants from multiple labs, Q4 down to Q1. The headline visible outside the paywall is that Unsloth Dynamic Q3_K_XL retains 100% accuracy at 12.8GB, which fits a 16GB card with room for context. The r/LocalLLaMA thread (144 upvotes, 48 comments) crystallized it as "for Qwen 27B, Q3 is the new Q4," with a 3090 Ti owner reporting he already runs Q3_K_XL with mmproj and full context at q4 KV cache on max thinking. Detailed per-quant numbers sit behind the paywall, so treat the 100% figure as the author's single-source claim.
↳ Follow the thread