Agents
Reasoning-channel prefill plus a trivial output prefix jailbreaks Gemini 3 Flash, DeepSeek V4 Flash and Claude Haiku 4.5 at up to 99%
Across 1,800 AdvBench cases, injecting malicious reasoning alone was inert at roughly 0% success, but pairing it with a short output prefix pushed attack success as high as 99% on some models, and contextual prefixes beat static ones. The attack needs an API that lets callers edit the reasoning scratchpad or prefill the response. Agent platforms that pass through prefill or thinking-block editing from untrusted callers should treat those fields as an injection surface.
Source
↳ Follow the thread