Fetching from the wire…
Public story · 2026-09-23 · high
A full-text meta-analysis of 259 agentic-security arXiv papers (Feb 2025 to Sep 2026) found 65.3% report no variance or repeated runs for their headline ASR.
Why now: This entered the 2026-09-23 research corpus from arXiv and has enough source evidence to become a public Rabbit Hole story.
A full-text meta-analysis of 259 agentic-security arXiv papers (Feb 2025 to Sep 2026) found 65.3% report no variance or repeated runs for their headline ASR. Only 30.9% disclose enough decoding detail to show whether the evaluation was stochastic, and 29.7% of LLM-judge papers check agreement with humans. On a 100-instance benchmark the minimum detectable ASR difference at conventional power is 18.2 points, so most published defense rankings with small gaps are noise.
Each link below shares sources, entities, or timing with this story.
The attack hides malicious intent across separate skills that only turn dangerous when they pass work to each other, and a fix cuts success to 22.5%.
ActBench ran 24,000 attack trajectories across 15 models and six harnesses; no harness pushed success below 73.7%.
The errors trace back to how the benchmark pairs pull requests with GitHub issues, not just to model quality.
It let 25,370 payments through and blocked only transfers to recipients the passport didn't list.
Standardizing how eight benchmarks score answers moved 9 of 10 models at least three ranks each.
A 4,181-problem study found confidence-based escalation beats a frozen router by 4.2 points while using 37% fewer tokens.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.