Tools
AMAP's LongHorizon-Harness Claims +28.9 Points on WeaveBench With a Manager/Executor/Auditor Split
AMAP-ML/LongHorizon-Harness (created 2026-08-04, 370 stars, MIT, v0.1.2 on 08-06 and v0.1.3 on 08-07, paper arXiv 2608.01964) separates long-running computer-use work into three roles with independently assignable models: Manager plans, Executor acts, Auditor verifies. Reported gains over baseline: WeaveBench 51.8% → 80.7%, Terminal-Bench 2.1 69.7% → 77.2%, and OSWorld 2.0 2.8% → 8.3% (a 3× jump off a brutally low floor). It integrates natively with Claude Code, Codex CLI, and OpenClaw via an `AgentAdapter`, defaulting to 30 rounds — the practical takeaway for builders is that a dedicated auditor role, not a bigger model, produced the delta.
Source
↳ Follow the thread