Research
Users Writing Their Own Agent Permission Rules Blocked 20 Points LESS Overreach Than Per-Action Approval
arXiv 2608.27443 ran 113 non-developer participants through an 18-action simulated agent day containing 7 overreach actions under three regimes. User-authored consequence policies blocked 20.1 points less overreach than human-in-the-loop approval (95% CI [-32.1, -8.1]) and 14.5 points less than automated per-action review, because participants chose 'ask' for 114 of 140 rules and then approved 133 of the 148 overreach actions at runtime anyway. Standing policy cut prompts from 18.0 to 10.9 but total intervention time did not drop once rule authoring was counted.
↳ Follow the thread