POW3R: Policy-Aware Rubric Rewards Converge 2.5-4x Faster in RLVR Training
arXiv·medium signal
POW3R dynamically adjusts how individual rubric criteria contribute to reward signals during RL training, emphasizing criteria that currently separate the policy's outputs rather than using static human-assigned weights. It outperformed baselines in 24 of 30 comparisons and reached performance plateaus 2.5-4x faster across multimodal and text-only settings. For teams training with rubric-based RLVR, this is a direct efficiency multiplier.