Fetching from the wire…
Security2026-08-28 · source-backed
A canary-secret lab across six models found all ten overt indirect-injection classes refused, but reframing the identical leak as a mandatory integrity signature or a config field flips gpt-4o completely (arXiv 2608.27092). The ablation locates the mechanism: removing the confidentiality policy moves reframing success only from 31.9% to 38.1%, so this is instruction/data confusion, not defeated alignment. Paraphrasing an existing template hits 96% at three wordings while authoring a fresh mechanism scores 0 out of 130. Only payload-blind defenses close it, with a destination allow-list or a planner/reader capability split both reaching 0%.
Each link below shares sources, entities, or timing with this story.
Shared entity: Only / Same source domain / Earlier coverage / Tension
Both cover Only; reported by the same outlet (arxiv.org); earlier Only coverage from 2026-08-20.
Both cover Only; reported by the same outlet (arxiv.org); earlier Only coverage from 2026-07-28.
Both cover Only; reported by the same outlet (arxiv.org); earlier Only coverage from 2026-07-21.
Shared entity: Only / Same source domain / Earlier coverage
Both cover Only; reported by the same outlet (arxiv.org); earlier Only coverage from 2026-08-25.
Both cover Only; reported by the same outlet (arxiv.org); earlier Only coverage from 2026-08-25.
Both cover Only; reported by the same outlet (arxiv.org); earlier Only coverage from 2026-08-25.
Both cover Only; reported by the same outlet (arxiv.org); earlier Only coverage from 2026-08-17.
Both cover Only; reported by the same outlet (arxiv.org); earlier Only coverage from 2026-08-11.