llama.cpp Gets Q2_0 CPU Quantization: a 1.7B Ternary Model Runs 2.4x Faster Than FP16 at 1/7th the Size With 99.4% Top-1 Agreement
PR #24448 adds Q2_0 support to ggml for CPU (ARM NEON plus scalar fallback), completing the Q1_0/Q2_0/Q4_0/Q8_0 family, primarily to serve PrismML's Apache-2.0 Ternary Bonsai models (1.7B/4B/8B). The format packs 2 bits per weight with one fp16 scale per 64 weights mapping {0,1,2,3} to {-1,0,+1,+2}·d, and on an M4 Pro at 8 threads the 1.7B goes from 3.20 GiB/48.70 t/s at F16 to 461.79 MiB/117.20 t/s, the 8B from 15.25 GiB/14.88 t/s to 2.15 GiB/30.14 t/s — with mean KLD of 0.00012–0.00020 and 99.31–99.39% identical top-1 tokens. Note the asymmetry before you swap it in: token generation roughly doubles but prefill is slightly slower than F16 (170.51 vs 200.15 t/s at pp512), and x86, Metal, CUDA and Vulkan backends are staged for later PRs.
↳ Follow the thread