Together AI: Distribution-Aware Speculative Decoding Cuts RL Post-Training Rollout Time by 50%
Together AI Blog·medium signal
Together AI's DAS framework addresses the rollout bottleneck in RL post-training, where 70% of training time is consumed generating responses before the next step can begin. Uses a training-free suffix tree drafter from recent rollouts that adapts without gradient updates, plus length-aware GPU load balancing. Achieves 50% rollout speedup with identical training curves on math and code tasks.