Research
HarnessDev Makes the Agent Harness the Unit of Evaluation, and Model-Built Harnesses Still Trail Human Ones on Code
HarnessDev shifts evaluation from task outputs to runnable infrastructure, asking an agent first to build a complete execution system from a minimal seed, then to revise its own harness using downstream execution feedback. Creation results cover six creator LLMs, four domains, and five downstream benchmarks totaling 2,207 instances with hidden evaluation tasks withheld. Generated harnesses remain substantially behind mature human-engineered references on code and on search and research while matching or exceeding them on writing and ML experimentation, and the Evolution gains are unstable, transfer only partially to held-out tasks, and depend strongly on which model executes the harness.
↳ Follow the thread