Fetching from the wire…
Research2026-09-03 · source-backed
Counter-GEO-Bench pairs 247 human-verified queries with matched information-preserving and information-distorting rewrites, scoring defenses across three victim LLMs. Granite Guardian, Llama Guard 3 and NeMo Self-Check Fact-Checking reduced attack success by at most 5.7% relative, and Granite Guardian's reduction wasn't statistically significant. Safety-taxonomy guardrails look for policy violations while GEO misinformation reads as fluent informational content. The paper's own baseline cut ASR 47.6% relative with near-zero utility loss, which says the problem is tractable with a purpose-built classifier and untouchable with a repurposed safety one.
Each link below shares sources, entities, or timing with this story.
Reflex-Guard combines jailbreak-aware preprocessing, compact sentence-transformer embeddings and seven binary classifiers trained on 30,568 samples, reporting 95.9% recall end-to-end against 255ms for Llama Guard 2 and 723ms for SafeDecoding, with 100% detection of GCG suffix...
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
The method synthesizes security skills offline from known attacks and recorded agent failures, injects them into the system prompt at session start, and leaves them active through the tool-use loop. Across six models on RedCode the default all-classes skill dropped malware-gen...
ADeptS-Bench tests seven models on paired benign and malicious GUI tasks and finds none clearing 80% task success while staying under 30% attack success. The ablation is the usable finding: removing the refusal tool raises attack success 21-23pp for tool-dependent models and 1...
Responsible Statecraft reported Aug 17 that the "Hanover Institute for Public Policy" is a front created by Piro, Inc. under a $900,000 contract from the Israeli Government Advertising Agency, subcontracted through Havas Media. Piro's own site markets the practice as "AI Story...
Sleeper Cell (2603.03371) — Two-stage attack embeds latent malicious behavior in fine-tuned tool-using LLMs. Poisoned models pass all benchmarks while harboring temporal trigger-activated harmful tool calls. Direct supply-chain risk for anyone using third-party LoRA adapters....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.