Self-State Attacks: 43 Concrete Operations That Compromise a Self-Hosted Agent Through Its Own Files
Self-hosted agents read and write their own memory and configuration files to function, so an attacker can compromise one entirely through legitimate OS system calls — no exploit required. The paper characterizes this threat class across Target, Mechanism, Granularity, and Temporal dimensions, instantiating a 23-cell attack matrix with 43 concrete operations against real self-state files, validated with live activity traces from a representative self-hosted agent. A layered defense (access control for instruction/config layers, workload-conditioned detection for the memory layer, periodic backup for recovery) covered most cells, but a residual attack surface remained structurally indistinguishable from legitimate activity at the OS level.
↳ Follow the thread