Research
HarnessLens Evolves Agent Harnesses With Behavior-Relevant Verification, Gaining 7.6-13.6% on Less Budget
arXiv 2608.27311 targets the cost of propose-and-verify harness tuning, where scoring every candidate on a fixed task set wastes rollouts on unrelated behaviors and lets aggregate scores hide specific regressions. HarnessLens jointly explores the task space and user-configurable components, derives candidate modifications from execution trajectories, and verifies each candidate only on behavior-relevant tasks through an attributable-evidence gate. Across three agent harnesses and four benchmarks it improves average held-out performance by 7.6-13.6% while consuming substantially less evaluation budget, with code at github.com/jhxu5214/HarnessLens.
↳ Follow the thread