Research
Repeat-After-Me Gets 47% Attack Success on GPT-5.5 With Black-Box Image Prompt Injection That Emits Real Tool Calls
arXiv 2609.04533 closes the gap between text and image prompt injection, which previously failed on frontier VLMs because harmful outputs require long, format-compliant strings like a parseable native tool call with exact function names and arguments. Repeat-After-Me exceeds 80% attack success on open-weight models including Qwen3.6-27B and 47% on commercial frontier VLMs including GPT-5.5, under a realistic setting where the benign user prompt is unrelated to the injected task. Injections optimized on one surrogate retain 43-46% of their success rate on two commercial victims, so surrogate-trained attacks transfer.
↳ Follow the thread