Agents
Cautious Bench measures the opposite failure: guardrails refusing authorized actions because the wording sounds scary
A 27 August paper argues over-safety in agent guardrails cannot be measured by harvesting real data, because boundary cases are rare and their labels are a function of an authorization policy rather than the action itself. Cautious Bench codesigns each sample with a stated policy and uses a build-time gate that re-derives every example so the label is a mechanical consequence of the policy, yielding 756 decidable benign/twin pairs rendered under three object-name types for 2,268 measured pairs, plus 40 undecidable pairs reported separately. Six guardrails from five providers are measured against it, giving builders the first reference for how often their action gate blocks legitimate work.
Source
↳ Follow the thread