Research
RewardFlow: Multi-Reward Langevin Dynamics Steers Diffusion Models at Inference Time Without Inversion
RewardFlow unifies multiple differentiable rewards (semantic alignment, perceptual fidelity, localized grounding, human preference) into a single inference-time framework for pretrained diffusion and flow-matching models. The inversion-free approach introduces a differentiable VQA-based reward providing fine-grained semantic supervision through language-vision reasoning. Practitioners can steer existing models toward specific quality objectives without retraining.
Source
↳ Follow the thread