Post-training on office workflows — with zero software-engineering tasks — improved SWE-Bench Pro pass@1 by 5.8 points
The authors define goal-directed execution as four repeated behaviors — selecting goals, constructing task-relevant state, maintaining fidelity to higher-level objectives, and verifying completion against the environment — and hypothesize that long-horizon post-training strengthens them independent of domain. Post-training Qwen3.5-122B-A10B on 363 long-horizon multi-tool tasks drawn purely from office workflows, containing no coding tasks at all, raised pass@1 on SWE-Bench Pro by 5.8 points, with matched trajectory analysis showing gains in all four behaviors in both office and repository settings. If it holds up, the practical implication is that long-horizon agentic competence is a transferable behavioral skill rather than domain knowledge — which changes what training data is worth collecting, and suggests coding-agent gains may come from unrelated long-horizon corpora.
↳ Follow the thread