Reddit
Four 2017 Tesla V100s Match a $6,000 RTX 5090 on Qwen3.8 NVFP4 Decode, on Silicon With No FP4 Support At All
A developer published v100-skinny, hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving, and reported four Volta V100s hitting 219.1 ± 5.9 tok/s decode against an RTX 5090 running the NInfer engine at 214.7 ± 9.2 tok/s, both 5/5 correct on AIME 2026 problem 1 across five seeds. The mechanism is the honest part: the V100 round takes 26.9ms versus 19.9ms (35% slower) but commits 5.89 tokens per round versus 4.27 (38% more), because QPN lets it run Qwen3.8's own MTP at depth 7 while NInfer caps at 5. The repo (68 stars, created 2026-08-11, last pushed 2026-08-19) is a concrete argument that speculative depth can buy back a whole hardware generation.
Source
↳ Follow the thread