Research
ParetoSlider Enables Continuous Multi-Objective Reward Control for Diffusion Model Post-Training
ParetoSlider replaces the standard single-scalar reward in RL post-training of diffusion models with a continuous control mechanism that lets users navigate the full Pareto front of multiple reward objectives at inference time. Instead of 'early scalarization' that bakes in fixed trade-offs, the approach trains once and allows dynamic adjustment of quality vs. diversity vs. aesthetic scores after deployment. Practical for any team fine-tuning image generators where different use cases need different quality-diversity trade-offs.
Source
↳ Follow the thread