Fetching from the wire…
Security2026-09-01 · source-backed
Across seven aligned models and three jailbreak attacks, holding the attack fixed and changing only a routine system prompt with nothing to do with safety shifted attack success by up to 56 points (arXiv 2608.30748). The increases showed up even for attacks tuned against the default configuration, and the swing tracks projections onto a refusal-related axis in hidden representations. Any red-team number measured in one prompt configuration does not transfer to the configuration you ship. Re-run your safety evals against your production system prompt, not the vendor's.
Each link below shares sources, entities, or timing with this story.
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
arXiv 2608.27141 proves a separation result: against an attack whose evidence is fragmented across iterations, any monitor whose safety state resets each trajectory has a true-positive rate equal to its false-positive rate, no matter how expressive it is. A monitor retaining c...
SecOPD fine-tunes a defense using token-level feedback during on-policy distillation rather than the sequence-level signal prior work used. Against PISmith adaptive injections on Qwen3.6-27B it reports 9.0% attack success where Meta-SecAlign, the previous state of the art, sit...
Reflex-Guard combines jailbreak-aware preprocessing, compact sentence-transformer embeddings and seven binary classifiers trained on 30,568 samples, reporting 95.9% recall end-to-end against 255ms for Llama Guard 2 and 723ms for SafeDecoding, with 100% detection of GCG suffix...
arXiv:2608.03070, submitted August 4 by Timm, Struppek, Gleave, Pelrine and 11 co-authors, composes 67 readily accessible static jailbreak techniques into an attack space and runs it against four frontier models over 360 goals spanning CBRNE and offensive cyber. A "universal j...
Willison and Jesse Vincent's Prime Radiant shipped a compact eval framework on a four-layer model: an eval contains tasks, each task runs against configs (model plus parameters like system prompt), each execution produces a run, and graders apply checks; string matches, custom...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.