Tools
llama.cpp lands Q4_K and Q6_K on Qualcomm Hexagon, unlocking Q4_K_M models on-device
PR #28994, merged 2026-09-16T16:00Z, adds Hexagon NPU support for the K-quants Q4_K (reusing the existing Q4_1 infrastructure) and Q6_K (new kernels), which together enable Q4_K_M models since those are typically a mix of the two. The author reports validation across multiple runs on IQ8 (VentunoQ) and IQ9 platforms. Q4_K_M is the default quant most people download from Hugging Face, so this closes the gap between what is published and what Hexagon-backed phones could actually run.
Source
↳ Follow the thread