Fetching from the wire…
Public story · 2026-07-27 · high
The discovered scenarios raised attack success by up to 18.2 points across six jailbreak methods, per the paper.
Why now: As of July 27, the paper was newly circulating and none of the three affected labs had said whether the discovered scenarios still work.
Researchers used a sparse autoencoder to find the internal concepts behind jailbreak scenarios, then reused them to raise attack success by up to 18.2 points, per a new paper.
That's a mechanism for jailbreaking, not a lucky prompt. The same discovered scenarios transferred to GPT-5, Claude-Haiku-4.5, and Gemini-3-Flash, three closed models built by different labs. Stacking scenarios together beat any single one and cut the turns an attacker needs.
The approach differs from most jailbreak research, which iterates on wording until something works. Concept2Scenario opens the model up instead. The authors trained a sparse autoencoder to represent internal activations as concepts, then traced which concepts suppress refusal when a scenario-wrapped prompt fires. They translated the highest-signal concepts back into natural-language scenarios and tested the results against six existing black-box jailbreak methods across three open models.
The paper doesn't say whether the three affected labs have patched the scenarios it discovered, or whether the technique scales past the concepts the authors found. That gap matters more than the topline number.
The bet worth making: prompt-level defenses tuned to catch specific wording will keep losing to attacks built from a model's own internal representations. An attacker only needs one lab's interpretability tooling to generate scenarios that travel to competitors' models. That's the shift here, from hand-written prompts toward attacks derived from a model's own internals.
Each link below shares sources, entities, or timing with this story.
Anthropic released Claude / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Anthropic released Claude); both cover Flash, Gemini, GPT; reported by the same outlet (arxiv.org).
Gemini competes with Claude / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Gemini competes with Claude); both cover Claude, Gemini, GPT; reported by the same outlet (arxiv.org).
Gemini competes with Claude / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Gemini competes with Claude); both cover Flash, Gemini, GPT; earlier Flash coverage from 2026-03-17.
Gemini competes with Claude / Shared entities / Same source domain / Earlier coverage
Linked by a graph relationship (Gemini competes with Claude); both cover Claude, Gemini, GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Gemini competes with Claude); both cover CLAUDE, Gemini, GPT; reported by the same outlet (arxiv.org).
Linked by a graph relationship (Gemini competes with Claude); both cover CLAUDE, Gemini, GPT; reported by the same outlet (arxiv.org).
Cursor supports Claude / Shared entities / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (Cursor supports Claude); both cover GPT, Haiku; reported by the same outlet (arxiv.org).
Cursor supports Claude / Shared entities / Earlier coverage
Linked by a graph relationship (Cursor supports Claude); both cover CLAUDE, Gemini, GPT; earlier CLAUDE coverage from 2026-07-20.