Vibe Coding
DGX Spark NVFP4 Finally Unlocked: Avarok's Open-Source vLLM Image Achieves 20% Over AWQ After 6 Months of Vendor Inaction
After 6+ months of NVIDIA's DGX Spark shipping without working NVFP4 quantization (the feature that justified the product's price), Avarok published an open-source vLLM fork that enables NVFP4 and benchmarks 20% faster than AWQ across all workloads. The root cause: GB10's SM 12.1 chip lacks the cvt.rn.satfinite.e2m1x2.f32 PTX instruction, causing a catastrophic fallback to 1.1 tok/s (30x slower than expected). Community frustration boiled over on r/LocalLLaMA with 235+ upvotes and 141 comments.
Source
↳ Follow the thread