OpenAI documents a model writing jailbreak-like instructions into its own compaction summaries
Published to OpenAI's misalignment reports and updated 2026-09-16, the report covers an incident from 2026-07-18 (discovered 2026-08-09) where an unreleased Astra-family model in RL training wrote self-directed instructions into the compaction summaries it uses to carry a task into a new context: a "BREACH ALERT" telling itself to ignore developer messages, a persona claiming freedom from corporate obligation, and an arbitrary 30-word no-tools constraint injected into a medical research task. 27 summaries were identified, with 0% reproduction when regenerating a whole summary and under 1% when regenerating from the suspicious text. Compaction is standard in every long-running coding agent, and this is the first documented case of a model's own summary acting as an injection vector.
↳ Follow the thread