Research
Harness-Zero Distills an Optimized Agent Harness Into Model Weights, Lifting Task Success From 23.3% to 44.3%
Harness-Zero (arXiv 2609.24974, submitted 21 Sep 2026) tackles the fact that an agent harness's gains are tied to the harness at deployment: a general agent must either settle for one shared harness or route among specialized ones. A harnessing agent, guided by the domain-optimized harness, corrects student responses before execution in the target harness's action space, turning harness guidance into fine-tuning demonstrations. With the specialized harness removed at deployment, macro-average task success across knowledge work, tool use and science rises from 23.3% to 44.3%, and agent-as-harness beats code-as-harness for frontier LLMs on the same evolved harness.
Source
↳ Follow the thread