Every trajectory-scoped agent monitor is provably useless against evidence split across loop iterations
A 27 August paper proves a separation result for autonomous agent loops: against an attack whose evidence is fragmented across iterations, any monitor whose safety state resets each trajectory has a true-positive rate equal to its false-positive rate, no matter how expressive it is, while a monitor that retains cross-iteration state separates them perfectly. The authors also show the obvious patch, a geometrically decaying risk score, fails because the cooling-off period a patient adversary must wait is a constant independent of the horizon N. Their LoopHarness restores a persistent non-decaying loop-level safety state and bounds expected unauthorized irreversible actions under mediated commits.
Source
↳ Follow the thread