Sources
DFlash 2 Quantizer Lands Support for Qwen 3.8 27B and Meta Muse Glimmer, With a llama.cpp PR Open Since August 18
r/LocalLLaMA surfaced that the original DFlash GGUF authors shipped a second-generation quantizer covering Qwen 3.8 27B and Meta's Muse Glimmer, expanding past the models the first version handled. The integration PR is real and verifiable: ggml-org/llama.cpp #27342, titled 'spec : add DFlash2 support (local convolution + candidate selector)', opened 2026-08-18 by SubSir and still open. Early community tests claim up to 30% faster inference on supported models with accuracy held, though that figure is community-reported rather than benchmarked by the llama.cpp maintainers, so treat it as provisional until the PR merges.
↳ Follow the thread