Research
Compacted Agent Memory Can Answer Today's Question and Still Break on Tomorrow's Update
arXiv 2609.20045 (17 Sep 2026) audits context compression with paired histories that share the same current answer, receive the same future update, and then require different answers. A deterministic frontier selector scored 96/96 strict reveal accuracy on DeepSeek but 82/96 on GLM, while a structured writer managed 56/96, and a record-level audit found 26 and 25 memories that were well formed but semantically wrong. Tombstone removal produced 16/16 exact replay failures in the targeted mechanism, and identifier renaming dropped frontier late-reference adequacy from 8/8 to 94/320 transformed instances, which is a direct warning for anyone compacting long agent sessions.
↳ Follow the thread