Fetching from the wire…
Security2026-09-15 · source-backed
PIDS-Bench evaluates seven prompt-injection detectors at fixed thresholds across in-distribution inputs, hard-benign prompts that mimic injection structure, obfuscated attacks and domain shifts, scoring false positives as a first-class axis (arXiv 2609.15017). A detector exceeding 0.98 F1 on its held-out split still misclassifies about a third of an externally-sourced benign subset restricted to security content. Across a full threshold sweep and five seeds, no internal detector hit F1 ≥ 0.95 and hard-benign FPR ≤ 0.10 simultaneously. If your product's users talk about security, this detector class will fight them.
Each link below shares sources, entities, or timing with this story.
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
Dahal and Xiong target injected documents that are individually benign but create false associations once aggregated, which is structurally invisible to any per-document filter (arXiv 2607.20437). TopoGuard builds a semantic similarity graph over the retrieved set and flags ma...
ADeptS-Bench tests seven models on paired benign and malicious GUI tasks and finds none clearing 80% task success while staying under 30% attack success. The ablation is the usable finding: removing the refusal tool raises attack success 21-23pp for tool-dependent models and 1...
Existing neuron-level defenses stay always-on and perturb every benign request (arXiv 2608.14392). Tripwire identifies safety-specific neurons through per-neuron hypothesis tests under false-discovery-rate control plus a utility-specificity filter, then clamps them to harmful-...
arXiv 2608.11878 replaces the handful of manually implemented injection-testing environments with an Environment Simulator, Attacker Agent, and User Simulator that generate executable stateful environments and discover viable injection points automatically. Injection timing an...
MetroLLM-Bench is 955 cases across six real metro systems of 37 to 414 stations, requiring structured tool calls and a machine-renderable terminal state. On the 238-case held-out split the 4B student scores 91.3 on Tier 1 against GPT-5.6's 90.6 and 90.0, at Q4_K_M. The gain ov...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.