Drifting Preference Optimization Enables Alignment of One-Step Image Generators Without Likelihoods
arXiv·medium signal
Solves a key obstacle in aligning one-step text-to-image generators: standard alignment methods rely on policy likelihoods and denoising trajectories that don't exist in single-forward-pass models. Drifting Preference Optimization introduces a new formulation that works without these dependencies. For teams deploying fast image generation (single-step distilled models), this unlocks RLHF-style alignment that was previously only possible with multi-step diffusion.