Reddit
POCKET-35B Ships GGUFs Down to 1.9 bpw for GPU-Free Inference — but the Model Card's Own Numbers Don't Match the 59 tok/s Claim
FINAL-Bench/VIDRAFT published POCKET-35B-GGUF, an Apache-2.0 repack of the Qwen3.5-35B-A3B MoE (256 experts, top-8 routing) pitched at running a 35B agentic model on a CPU-only PC or a phone via stock llama.cpp. Quants run from Q4_K_M (21.2GB) to IQ1_M (8.24GB, 1.9 bpw), with card-reported GPQA-Diamond of 68.7% at Q4_K_M and HellaSwag 61.0% at IQ1_M; 6,551 downloads in the last month. Worth flagging the gap: the r/LocalLLaMA post advertises 59 tok/s on CPU, while the card itself cites 27.0 tok/s on a 16-thread Xeon (IQ1_M), 19.5 tok/s on M3 Pro CPU at Q2_K, and 13.8 tok/s on an 8-thread M3 Pro. Benchmark before you plan around the headline number.
↳ Follow the thread