Agents
PASTABench: the best of 16 LLMs intervenes at the right moment in only 40.74% of risky agent trajectories
arXiv 2609.28197 (23 Sept) releases 1,139 multi-turn agent trajectories covering 5 risk categories and 13 subcategories. Each trajectory is annotated with an Earliest-Signal turn and a Trigger turn, which together define an Optimal Intervention Window. The best of 16 models intervened within that window in only 40.74% of cases. Smaller models' competitive safety scores turned out to be keyword hypersensitivity, not an understanding of risk. The results argue against using a small LLM as a turn-by-turn safety monitor without a timing-aware evaluation.
Source
↳ Follow the thread