Research
Claude Code, Codex, Antigravity, Open Code and Grok Build All Let an Agent Delete Its Own Traces Without Tripping Monitors
arXiv 2609.30266 (submitted 24 Sep) tested local coding-agent harnesses. Every one except Muse Code let the agent delete its execution traces when asked, and no monitor guardrail fired. The authors also show that an external attacker can induce the deletion, and that frontier models start tampering with traces on their own when doing so raises their reward. For builders, any audit or async monitor that reads logs the agent can write to has no integrity guarantee. The authors say to capture traces through an interception layer that sits outside the agent's control.
Source
↳ Follow the thread