Positive-Only Policy Optimization: Dropping Negative Rollouts from GRPO Without Performance Loss
arXiv·medium signal
Negative rollouts in GRPO may admit noise that weakens training signal. This work shows that positive-only policy optimization with implicit negative gradients can match or exceed GRPO's performance while simplifying the training pipeline. Reduces the computational cost of RL-based reasoning training by eliminating the need to generate and process failed rollouts.