FP4 Pretraining With E5M3 Block Scales Beats NVIDIA's Transformer Engine Recipe and Runs 21% Faster
Stable 4-bit pretraining is hard because the E2M1 payload covers only a narrow magnitude range; NVIDIA's Transformer Engine recipe compensates with current-tensor scaling, a randomized Hadamard transform and BF16 final layers, all of which add work outside the FP4 matmuls. This work pairs E2M1 payloads with unsigned E5M3 block scales whose wider range permits periodic tensor scaling, applies selective stochastic rounding only to backward gradients, drops the Hadamard transform entirely, and uses FP4 in every eligible internal linear. Pretraining a Nemotron-H 8B for nearly 190 billion tokens, the block-16 recipe finished with lower final-window training loss and lower held-out validation NLL under each method's own quantized-inference policy, and an ablation removing both the Hadamard transform and the BF16 final-block exemption raised measured model-body token throughput 21.2%.
↳ Follow the thread