Ternary Bonsai 2 27B ships a true 1.72-bit quantization of Qwen3.8-27B at 5.95 GB, and hit 405,000 downloads in under two days
Prism ML released Bonsai 2 27B on 2026-09-16 as GGUF and MLX builds that store embeddings, attention projections, MLP projections and the LM head as ternary {-1,0,+1} weights with FP16 group scaling, cutting a ~54 GB FP16 model to 5.95 GB (PTQ1_0 packing) while keeping the 262K context and hybrid-attention backbone. The card claims 98.2% of FP16 intelligence retained, an 84.78 average across 14 thinking-mode benchmarks against 72.59 for a conventional IQ2_XXS build, agentic tool calling at 74.92, and ~47 tok/s on an M5 Max laptop. The GGUF repo shows 405,609 downloads and 612 likes two days after creation, but a same-day r/LocalLLaMA counter-thread ('Ternary Bonsai is a headless chicken') reports the model failing a single 3D-scene coding prompt, so treat the 98.2% figure as vendor-measured until third parties reproduce it.
↳ Follow the thread