Fetching from the wire…
Public story · 2026-09-24 · high
PASTABench found some small models scored well by matching keywords, not by understanding risk, across 1,139 agent trajectories.
Why now: New coverage from September 24, 2026 evidence.
PASTABench tested 16 AI models on catching risky agent conversations before they escalate. The best one hit the right moment only 40.74% of the time.
That gap matters for anyone building an agent overseer. Miss the window and the risky action already happened. Flag too early and users learn to ignore the warning.
The PASTABench paper built 1,139 multi-turn agent trajectories across 5 risk categories. Each trajectory got two human-annotated markers.
An Earliest-Signal turn marks where a risk first becomes visible. A Trigger turn marks the point past which intervening is too late. The space between them is the Optimal Intervention Window, and models earned credit only for flagging the conversation inside it.
Small models turned out to be the surprise. Several posted safety scores competitive with much larger models, but the paper attributes this to keyword hypersensitivity rather than real risk understanding. Those models flagged individual trigger words. They didn't track where the conversation was heading.
The paper doesn't say whether the gap closes with more training or a different architecture, and it doesn't name which of the 16 models scored highest. What it does establish is a bar. Even the best model missed the intervention window more than half the time.
Each link below shares sources, entities, or timing with this story.
Material Discovery Bench measures LLM progress on discovering thermally conductive dielectrics for 3D chip packaging, built on the AI Security Institute's open-source Inspect framework, with models given web search, Python/bash sandboxes, and property-computation ML tools (Dis...
Agents get only a CWE description and terminal access, and have to find the implementing files across 500 real vulnerabilities from 290 repositories, six package ecosystems and 147 CWE categories (arXiv 2609.15939). Across 27 language models and four static-analysis tools on a...
arXiv 2608.24358 switched models mid-run on long coding tasks using cheap/expensive pairs from the Claude and GPT families. Full-trajectory escalation from weak to strong recovers under half the gap while costing a substantial premium, which the authors call the handoff tax. D...
The proposed elements include agent memory and memory access records, actual versus potential autonomy level, and tool usage, none of which existing AI incident frameworks capture. The experts also flagged the reporting pipeline as its own attack surface, since incident record...
Researchers loaded five systems with a revoked policy and its replacement, then measured retrieval and downstream action across nine policy scenarios, nine models and six defense conditions. Wherever the revocation label was visible to the retrieval layer, the revoked fact cam...
DSEffi-Bench covers 1,000 instances across 10+ libraries with stress-testing harnesses and human-validated references, evaluated on 16 models (arXiv 2608.30248). GPT-5.4 leads correctness at 66.9% Pass but its 71.7% efficiency score barely beats GPT-5.4-mini's 71.6% despite so...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.