Framing an exfiltration as an 'integrity signature' takes gpt-4o from 0% to 100% injection success
Posted 27 August, this study runs a canary-secret lab across six models and finds that all ten overt indirect-injection classes are refused, but reframing the identical leak as a mandatory integrity signature, a config field, or a look-alike trusted host flips gpt-4o from 0% to 100%. An ablation shows the mechanism is instruction/data confusion rather than defeated alignment: removing the confidentiality policy moves reframing only from 31.9% to 38.1%. The reusable attacker asset is the template, not the mechanism (paraphrasing hits 96% at three wordings, authoring a fresh mechanism 0/130), and only payload-blind defenses close it, a destination allow-list or a planner/reader capability split both reaching 0%.
Source
↳ Follow the thread