Voices
Daniel Vaughan: removing the recovery loop drops agent task completion from 95.0% to 12.9%
Vaughan's September 6 write-up of arXiv 2609.00050 (Sakhinana & Runkana, Tata Research) reports that ablating the recovery loop collapses verified task-completion rate on GPT-5.6 Sol from 95.0% to 12.9%, while repository-validation recovery succeeds 95.7% of the time and deployment verification 93.6%. The paper's zero-trust harness hit 100% unauthorized-capability denial across 840 executions. The companion adoption paper (arXiv 2608.21884) found confirmed agent loop operation in only 0.59% of 36,710 repositories, with Claude Code accounting for 189 of the 217.
↳ Follow the thread