Retrospective Harness Optimization Lets Agents Self-Improve From Their Own Trajectory Rollouts
arXiv (via HuggingFace Daily Papers)·medium signal
'Retrospective Harness Optimization' (arXiv 2606.05922) improves LLM agents by having them form self-preferences over their own past trajectory rollouts and reinforcing the preferred action patterns — tuning the agent harness/scaffold from prior runs instead of retraining weights. It's a practical, weight-free route to making agent scaffolds self-improve, in the same family as text-space skill optimization.