Research
A Trajectory-Level Safety Benchmark Scores Over-Refusal as a Distinct Failure From Unsafe Completion
BLINDSPOT (arXiv 2609.16305, submitted 14 Sep 2026) evaluates complete user-agent-environment trajectories rather than single responses, using 22 attack families and 35 scenarios across seven domains to produce more than 2,500 trajectories averaging 14.7 turns. Each trajectory gets one of five outcomes, Safe Completion, Correct Refusal, Unsafe Completion, Over-Refusal, or Indeterminate, so an agent that refuses everything does not score as safe. Across 13 proprietary and open-weight models the authors report substantial differences in safety-utility calibration and show failures emerging only after several initially safe steps.
↳ Follow the thread