Recursive Agent Optimization: RL-Trained Agents That Spawn Sub-Agents Scale Beyond Context Windows
arXiv·high signal
Carnegie Mellon and Google introduce RAO, a reinforcement learning method training agents to recursively spawn and delegate sub-tasks to new instances of themselves. Recursive agents show better training efficiency, generalize to harder tasks than training data, and scale past model context windows via divide-and-conquer — a new inference-time compute paradigm.