Proposes on-policy distillation for flow matching models that overcomes the reward sparsity problem under multi-task alignment — where off-policy methods degrade because the student distribution diverges from the teacher. Enables simultaneous optimization for aesthetic quality, prompt adherence, and safety without the quality collapse seen in prior approaches. Practical for teams fine-tuning open-source image generators.