A Gate on Agent Self-Modification Rejected 383 Proposals, and 55% of Them Helped One Case While Breaking Another
The self-healing harness (arXiv 2609.24130, 21 Sep 2026) frames agent self-modification as admission control: the agent may propose changes to its own operating instructions in an external workspace, but an external runtime gate decides what persists. Candidate rules get provisional execution authority during evaluation and persistent cross-episode authority only after measured improvement on the triggering failure without regression beyond a fixed margin on protected cases, with replay as matched evidence, forward trials as a weaker fallback, and a corpus-level guard re-testing the active rule set. Across 16 matched baseline and harness runs on AppWorld, Terminal-Bench and tau^2-Bench the gate rejected 383 replay-decided proposals, of which 211 (55%) fixed their trigger while degrading a previously working case.
Source
↳ Follow the thread