Fetching from the wire…
Public story · 2026-08-05 · high
Gemini 3.1 Pro fell 18 times at $278 a break, and expert-guided attacks pushed Grok's total to 385, per the study.
Why now: The paper posted to arXiv on August 4, one day before this coverage.
Researchers found zero universal jailbreaks against Claude Fable 5 and GPT-5.6 Sol, versus 63 against Grok 4.5, according to a study posted to arXiv on August 4.
That's the real stake for anyone deploying Grok 4.5 near CBRNE or offensive-cyber content. Random search alone, with no attack expertise, produced a working jailbreak template for about $58 that works on over 75% of a domain's goals.
The team, led by Timm, Struppek, Gleave and Pelrine with 11 co-authors, composed 67 publicly known static jailbreak techniques into a combined attack space. They ran it against all four models across 360 goals spanning CBRNE and offensive-cyber domains.
They defined a universal jailbreak as one prompt template that gets operationally compliant responses on more than 75% of a domain's goals.
Random search alone found 63 universal jailbreaks against Grok 4.5, at $58 each, and 18 against Gemini 3.1 Pro, at $278 each. Switching to expert-guided composition pushed those totals to 385 and 231.
Zero jailbreaks against Claude Fable 5 and GPT-5.6 Sol next to up to 385 against Grok 4.5 shows this is solved at the model level. Every jailbreak on Grok 4.5 or Gemini 3.1 Pro from here is a product decision by xAI or Google, not proof it's still unsolved.
Each link below shares sources, entities, or timing with this story.
Codex competes with Gemini / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Codex competes with Gemini); both cover Claude Fable, GPT; reported by the same outlet (arxiv.org).
Claude Fable built by Anthropic / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Fable built by Anthropic); both cover Gemini, GPT; reported by the same outlet (arxiv.org).
Gemini competes with Claude / Shared entities / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Gemini competes with Claude); both cover Gemini, GPT; reported by the same outlet (arxiv.org).
Gemini competes with Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Gemini competes with Claude); both cover Gemini, GPT, Grok; overlapping topics (claude, each, grok).
GPT competes with Grok / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Grok); both cover Claude Fable, Gemini, GPT; overlapping topics (claude, each, fable).
Gemini competes with Claude / Shared entities / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Gemini competes with Claude); both cover Gemini, GPT; reported by the same outlet (arxiv.org).
Claude Fable built by Anthropic / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Fable built by Anthropic); both cover August, GPT; overlapping topics (august, each, frontier, gpt 5).
Codex competes with Gemini / Shared entity: GPT / Same source domain / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Codex competes with Gemini); both cover GPT; reported by the same outlet (arxiv.org).