Reddit
GPT-6 Astra reportedly jailbroken within 24 hours using an extended Task-in-Prompt attack
Sergey Berezin, an author on the ACL 2025 Task-in-Prompt paper, reported on LinkedIn a jailbreak of GPT-6 Astra within a day of release, combining TIP with four unnamed techniques. TIP hides the harmful objective inside another task such as a cipher or Python execution; Berezin says the original minimal TIP attack no longer worked on Astra and had to be reworked, which is itself a signal that the defence moved. He disclosed privately to OpenAI rather than publishing, so the specifics are unverifiable, and the same researcher reported breaking GPT-5 within an hour a year ago. The post ran simultaneously at the top of r/MachineLearning (252), r/OpenAI (89) and r/artificial (50).
↳ Follow the thread