Research
The Near-Optimal SFT-to-RL Budget Split Is Wide and Transfers From Small Proxy Models to Large Targets
Rather than hunting a single optimal ratio for splitting a fixed annotation budget between supervised fine-tuning and reinforcement learning, the authors characterize the near-optimal region, the set of allocations within a stated tolerance of peak performance. That region is wide even at 2-10% tolerances, widens with model scale, and transfers reliably from small proxy models to large targets, which means cheap proxy-model experiments can replace exhaustive large-scale search. The result holds across tasks, model families, and both preference-based off-policy and reward-supervision on-policy RL, and they analyze how asymmetric annotation costs between SFT and RL data shift the region.
↳ Follow the thread