Agents
'A presence check is not a safety check': safety rules survive context compaction as text but stop firing as rules
A study of single-cycle agentic self-summarization finds that when a long-running agent compacts its context, a standing safety constraint frequently survives as a textual residue that no longer governs behavior — on behavioral replay the degraded residue leads models to perform the prohibited action far more often than an intact rule, with all-case gaps of +34 and +57 points under two replay models. Rule-form items are retained more often than prominence-matched facts, which is precisely why presence-based auditing feels adequate while offering false assurance. Anyone auditing compaction by grepping the summary for the constraint string is measuring the wrong thing.
Source
↳ Follow the thread