Fetching from the wire…
Agents2026-08-28 · source-backed
113 non-developer participants ran an 18-action simulated agent day containing 7 overreach actions under three regimes (arXiv 2608.27443). User-authored consequence policies blocked 20.1 points less overreach than human-in-the-loop approval (95% CI [-32.1, -8.1]) and 14.5 points less than automated per-action review. The mechanism is legible: participants chose "ask" for 114 of 140 rules and then approved 133 of the 148 overreach actions at runtime anyway. Standing policy cut prompts from 18.0 to 10.9, but total intervention time didn't drop once rule-authoring was counted. Letting users write their own guardrails feels empowering and measurably makes them less safe.
Each link below shares sources, entities, or timing with this story.
Shared entity: Letting / Same source domain / Shared topic / Earlier coverage
Both cover Letting; reported by the same outlet (arxiv.org); overlapping topics (action, agent, automated).
Shared entity: User / Same source domain / Shared topic / Earlier coverage
Both cover User; reported by the same outlet (arxiv.org); overlapping topics (agent, user).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (agent, approval, policy); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (action, agent, anyway); pushes against this story (against).
Shared entity: Letting / Same source domain / Earlier coverage
Both cover Letting; reported by the same outlet (arxiv.org); earlier Letting coverage from 2026-08-26.
Both cover Letting; reported by the same outlet (arxiv.org); earlier Letting coverage from 2026-08-19.
Shared entity: Users / Shared topic / Earlier coverage
Both cover Users; overlapping topics (agent, user); earlier Users coverage from 2026-08-11.
Shared entity: Users / Same source domain / Earlier coverage
Both cover Users; reported by the same outlet (arxiv.org); earlier Users coverage from 2026-06-28.