Research
DREvo Stabilizes Agent Harness Self-Evolution, Gaining 16.2% on Reasoning and 14.2% on Agentic Tasks
DREvo (arXiv 2607.26722, 2026-07-29) targets a real problem in self-improving agent loops: accumulated trial history does not reliably translate into search guidance, so harness performance oscillates wildly across evolution iterations under a limited budget. It adds function-level evidence anchoring, state-dependent evidence recalibration (deciding whether past experience is still valid for the current harness), and role-conditioned search intent distillation. Under constrained evolution budgets it produces smoother trajectories, takes the highest accuracy on all five benchmarks, and averages 16.2% and 14.2% gains over baselines on domain-reasoning and agentic tasks respectively.
↳ Follow the thread