Fetching from the wire…
Infra2026-08-20 · source-backed
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The mechanism is the honest part: the V100 round takes 26.9ms vs 19.9ms (35% slower) but commits 5.89 tokens per round vs 4.27 (38% more), because the setup runs Qwen3.8's own MTP at depth 7 while NInfer caps at 5. Speculative depth bought back a hardware generation.
Each link below shares sources, entities, or timing with this story.
Shared entities / Same source domain / Earlier coverage
Both cover FP4, LocalLLaMA, NVFP4, Qwen3; reported by the same outlet (reddit.com); earlier FP4 coverage from 2026-08-12.
Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Both cover AIME, LocalLLaMA, Qwen3; reported by the same outlet (reddit.com); overlapping topics (aime, qwen3).
Shared entities / Same source domain / Earlier coverage / Tension
Both cover LocalLLaMA, Qwen3, RTX; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-04-23.
Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Both cover LocalLLaMA, RTX; reported by the same outlet (reddit.com); overlapping topics (against, caps).
Shared entities / Shared topic / Earlier coverage
Both cover MTP, NVFP4, Qwen3; overlapping topics (nvfp4, speculative); earlier MTP coverage from 2026-08-18.
Shared entities / Same source domain / Earlier coverage
Both cover LocalLLaMA, NVFP4, Qwen3; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-08-13.
Both cover LocalLLaMA, Qwen3, RTX; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-08-10.
Both cover LocalLLaMA, Qwen3, Speculative; reported by the same outlet (reddit.com); earlier LocalLLaMA coverage from 2026-07-27.