Flow-GRPO's Scattered Instabilities Trace to One Measurable Per-Step Quantity
RL on flow-matching image generators is unstable in a way specific to multi-step denoising: importance ratios drift below one, disperse, clip at different rates and leave fewer usable samples late in training, and prior work patched each symptom with its own hand-tuned stabilizer. This paper shows all of them follow from a single per-step quantity it calls path variance, determined exactly by the sampler's Gaussian transition kernel and cheap to estimate during training. λ-Controlled GRPO calibrates importance-ratio behavior from that predicted law instead of noisy empirical statistics and allocates gradient effort across denoising steps by predicted cost, with the two governing scales fixed by standard policy choices rather than added as free hyperparameters.
↳ Follow the thread