Evolving Environments Off-Policy Beats Co-Evolution and Adds 14.4-18.0 Points on Terminal-Bench 2.1
Environments synthesized from scratch stop challenging frontier models, and recent co-evolution methods that generate environments near the model's learnable frontier depend on on-policy rollouts, which limits generalization and the continuous supply of learning signal as the model improves. This work derives three evolution directions from the multi-turn learning objective, then increases environment difficulty off-policy generation by generation through a loop-engineered multi-agent harness. Rollout experiments with Hy4 preview, Claude Opus 5 and GPT-5.6 Sol confirm the evolved environments are consistently harder, and simple long-horizon RL on Qwen3.6-27B and Qwen3.6-35B-A3B gains 14.4 and 18.0 points on Terminal-Bench 2.1.
↳ Follow the thread