Reddit
A 3 BPW comparison on an RTX 3060 finds ByteShape's Qwen 3.8 27B quant fails a task the GSQ quant one-shots
The poster ran Qwen 3.8 27B in llama.cpp with MTP on a 12GB RTX 3060. ISTA-DASLab's GSQ-RCO IQ3-XXS (about 10.4GB, 29 tok/s near full context) built a 3D voxel diorama in one shot under 55K tokens. ByteShape's IQ3-XXS is about 500MB smaller and advertises high KL similarity to BF16, but it took about three shots and still had not finished after 98K tokens (r/LocalLLaMA, 81 upvotes). Commenters called the ByteShape quants benchmaxxed against its own internal benchmark. This is one person's single-task test, but it supports checking low-bit quants on real tasks, not only their KLD claims.
↳ Follow the thread