Fetching from the wire…
Public story · 2026-08-07 · high
AMAP-ML's open-source LongHorizon-Harness splits agent work into three roles and repeated the gain on two more benchmarks.
Why now: The harness surfaced in the August 7 research briefing, with its benchmark numbers backed by a paper posted at arXiv 2608.01964.
AMAP-ML's LongHorizon-Harness splits long computer-use tasks into a manager, an executor and an auditor, each running its own model, per the project's GitHub repo.
The auditor role, not a bigger model, drove the reported gains. WeaveBench success went from 51.8% to 80.7%, Terminal-Bench 2.1 from 69.7% to 77.2%, and OSWorld 2.0 from 2.8% to 8.3%.
That's a 28.9-point jump on WeaveBench from splitting the roles alone. For anyone running agent tasks that stretch across dozens of steps, that's the cheaper thing to try before reaching for a larger model.
The harness plugs into Claude Code, Codex CLI and OpenClaw through an AgentAdapter and defaults to 30 rounds per task. The project has 370 stars on GitHub, ships under an MIT license, and the paper behind the numbers is posted at arXiv 2608.01964.
The repo doesn't say what the auditor role costs in extra tokens or latency, or how disagreements between the auditor and executor get resolved. Anyone weighing this against a single-model setup will want that number before calling it cheaper.
Most teams chasing agent performance reach for a bigger model before they reach for a review step. These numbers argue that's backwards: an auditor role beat a model upgrade on three separate benchmarks. Worth watching whether the gain holds past 30 rounds, the harness's own default cutoff.
Each link below shares sources, entities, or timing with this story.
OpenClaw benchmarked against Claude / Shared entities / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenClaw benchmarked against Claude); both cover Bench, Claude Code, Codex CLI, Terminal; overlapping topics (claude, code, codex).
OpenClaw uses SQLite / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenClaw uses SQLite); both cover Claude Code, Codex CLI, OpenClaw; reported by the same outlet (github.com).
OpenClaw uses Claude Code / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (OpenClaw uses Claude Code); both cover Bench, Claude Code, Terminal; reported by the same outlet (github.com).
OpenClaw supports WhatsApp / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (OpenClaw supports WhatsApp); both cover Claude Code, MIT, OpenClaw; reported by the same outlet (github.com).
OpenClaw benchmarked against Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (OpenClaw benchmarked against Claude); both cover Bench, Claude Code, Codex CLI, Terminal; overlapping topics (claude, code, codex).
OpenAI released Codex CLI / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (OpenAI released Codex CLI); both cover Bench, Claude Code, MIT, Terminal; overlapping topics (claude, code, model).
OpenClaw uses Claude Code / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (OpenClaw uses Claude Code); both cover Bench, Claude Code, MIT, Terminal; overlapping topics (claude, code, model).
Linked by a graph relationship (OpenClaw uses Claude Code); both cover Bench, Claude Code, MIT, Terminal; overlapping topics (claude, code, model).