Controlled Study of Long-Horizon Planning Finds Explicit World Models Generalize and Suboptimal Trajectories Actively Hurt
'The Physics of Multi-Turn Long-Horizon Planning' (arXiv 2607.24720, July 27, from Chinese Academy of Sciences authors) builds a unified controlled multi-turn environment to isolate where planning ability actually comes from, tracing it across three stages: acquisition during pre-training (data format and quality), shaping via post-training with GRPO and on-policy distillation, and integration through multi-teacher on-policy distillation (MOPD). Key findings: pre-training data containing explicit world models improves generalization, suboptimal trajectories degrade performance specifically over long horizons, and cross-environment learning works when planning patterns are compatible but causes interference when they conflict. For teams curating agent training data, the 'suboptimal trajectories are worse than fewer trajectories' result is the actionable one.
↳ Follow the thread