Research
Disaggregated Quantization Trains a Separate NVFP4 Prefill Checkpoint and Lifts 1-Bit Qwen3.8-27B by 32.5 Points on MMLU-Pro
Panferov et al. quantize the prefill phase and the decode phase differently. Decode gets compact weights to cut memory traffic, and prefill gets its own compute-native low-precision weights. On the released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller raised 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without touching the decode checkpoint. Streaming the extra prefill weights from SSD gave a 1.78x time-to-first-token speedup at 8K prompts in llama.cpp. If you run local models, prefill and decode no longer have to share one quantization compromise.
Source
↳ Follow the thread