Reddit
The hyped Bonsai ternary quant of Qwen3.8 27B finished the same tasks but burned 3.10x more output tokens and took 3.02x longer
A builder ran Qwen3.8 27B IQ3_XXS (10.18 GiB) against Ternary-Bonsai-2-27B-PQ2_0 (6.42 GiB) on four UI generation tasks on one 16 GB card. Both completed 4/4 tasks and fixed 20/20 assertions, but agent wall time was 8:00 versus 24:09 and output tokens 27,197 versus 84,176, with weighted decode 83.59 versus 64.91 tok/s and speculative acceptance 65.22% (MTP) versus 39.76% (modified N-gram). The lesson for anyone squeezing a big model into small VRAM: the smaller file's quality gap is narrow, but the token-count blowup makes the wall-clock cost far worse than the file size suggests.
↳ Follow the thread