Claimed: 366 tok/s Single-Stream on Qwen3.6 27B in NVFP4 Running on Volta-Era V100s
r/LocalLLaMA (106up/90c)·low signal
An r/LocalLLaMA poster reports 366 tokens/sec single-stream on Qwen3.6 27B quantized to NVFP4 running on V100s, following up on an earlier post claiming 1,000 tok/s aggregate generation on the same hardware. The thread has 106 upvotes and 90 comments — a 0.85 comment-to-score ratio indicating heavy technical scrutiny, which is warranted: V100 is Volta and has no native FP4 path, so the result depends entirely on the emulation approach used. Single-source and unverified by any benchmark harness, but if it holds it materially changes the economics of used-V100 clusters.