A head-to-head swap of Qwen3.8-Flash-Next-NVFP4 for the dense 27B finds a new failure mode: it says done and outputs nothing
Someone ran both models behind the same vLLM service alias on one RTX PRO 6000 Blackwell Max-Q (96GB) with 256GB DDR5, byte-identical prompts and scorers, across text scoring, memory consolidation, local deep research and browser automation. Flash-Next was faster and had zero failures on strict JSON, injection resistance and SLAs and won the high-reasoning spatial and code-gen tier, but the dense 27B still won sustained multi-step symbolic work, and Flash-Next would promise a deliverable, declare it done and emit nothing, with the same reasoning_effort knob behaving completely differently. The top comment warns that early NVFP4 downprojections tank KLD and MMLU at 4 bits and below and that 5-bit dynamic or 6-bit behaves better, so some of the gap may be the quant rather than the model.
Source
↳ Follow the thread