Reddit
A Full NVFP4 vLLM Recipe Fits Qwen3.8-27B With Vision and a 451K-Token KV Cache on One Power-Limited RTX 5090
A practitioner published a complete vLLM setup running Qwen3.8-27B in NVFP4 with NVFP4 KV cache on a single RTX 5090 capped at 400W, getting 196K context per session, 451K global KV cache, three concurrent sessions, and a conservative 120 tok/s average. Their llama-benchy 0.4.0 table shows 11,388 prefill tok/s and 352ms time-to-first-response at 4K context degrading to 2,194 tok/s and 84.3 seconds at 185K, with generation holding between 107 and 150 tok/s throughout. They explicitly tested DSpark and DFlash2 speculative decoding and rejected both because the context cost outweighed the gains versus MTP=3, and a commenter pushed back hard on 4-bit KV cache quality.
Source
↳ Follow the thread