Alibaba's AMAP Lab Published a Long-Horizon Harness That Lifts OSWorld 2.0 Completion 3x by Refusing Unverified Progress
AMAP-ML released LongHorizon-Harness on August 4 (463 stars, v0.1.3 shipped August 7) built on a strict Manager/Executor/Auditor split where, in its own words, 'only results that pass independent verification enter persistent task state.' The reported gains are large: WeaveBench 51.8% to 80.7% completion, OSWorld 2.0 2.8% to 8.3% full completion, and Terminal-Bench 2.1 69.7% to 77.2% success using 24% fewer tokens. It wraps Codex and Claude Code today via an AgentAdapter protocol, defaults to 30 rounds, and persists goals, verified progress, audit reports, and role trajectories in isolated run directories — the practical takeaway is that the auditor role, not a bigger context window, is what stops multi-hour runs from drifting.
Source
↳ Follow the thread