Vibe Coding
Quantization Choice Flips Up to 5.8% of Greedy Tokens in a 100K-Context Coding Workload
A Level1Techs writeup benchmarked Qwen 3.6-27B and 3.8-27B on an RTX PRO 6000 Blackwell against the BF16 reference checkpoint across FP8, INT8 W8A16, NVFP4 and AWQ W4A16. Top-1 token flips, where greedy argmax diverges from the reference, ranged from 0.717% to 5.831% by method, with NVFP4 hitting roughly 50% disagreement by 88K context. Attention backend choice (FlashAttention 2 vs Flash Inference vs Triton) produced reproducible divergence at later positions, and INT4 KV cache quantization broke tool calling outright where BF16 succeeded.
↳ Follow the thread