SkillGuard Shrinks Agent Authority After Untrusted Data Arrives, With Zero Extra Model Calls
Rather than classifying untrusted content or approving individual operations, SkillGuard treats untrusted data entering agent state as contamination and restricts future capabilities so the state cannot reach deployer-defined forbidden states, using a Skill Impact Graph, steerability signatures and an inline reference monitor. Across four AgentDojo suites with Gemini 2.5 Flash and Llama3.3-70B backends, it eliminates attack success on three of four suites for both models and cuts Slack to 4.8% and 14.3%, beating Spotlighting, CaMeL and AttriGuard. Fractional-flow restriction preserves substantially more capability than binary cutoff at equal attack success, and the whole layer adds no model calls or token overhead.
↳ Follow the thread