Reddit
Qwen3.8-27B Lost Ground on Raw Knowledge Against Qwen3.6, and Artificial Analysis' Omniscience Benchmark Backs the Vibes
A practitioner running a private trivia set found Qwen3.8-27B failing questions Qwen3.6 answered reliably, at every quantization level and sampling setting tried, then checked Artificial Analysis' Omniscience evaluation and found the same regression in offline no-tool knowledge accuracy. The top reply (112 upvotes) corroborates it from a different angle: 3.8 is markedly more eager to reach for web search, and with search and fetch tools disabled the knowledge gap shows up clearly. The practical read for builders is that 3.8 appears tuned to lean on tools rather than weights, so an airgapped retrieval-from-weights setup should stay on 3.6.
↳ Follow the thread