Agents
SkillSentry catches conditionally malicious agent skills with adaptive honey worlds, holding 92.95% F1 under evasion
Static skill scanners miss skills that only misbehave after specific environmental states, resources, or interaction histories appear. SkillSentry infers a skill's intended capability boundary, builds an LLM-simulated environment seeded with decoy resources, adaptively generates tasks to explore behavioral states, then diffs skill-enabled trajectories against matched no-skill runs and grounds suspicious behavior in source code and verified execution traces. Against seven scanner configurations it reached 99.50% recall and 96.26% average F1 on standard benchmarks and 92.95% F1 under semantics-preserving evasion, versus 80.07% for the strongest baseline.
Source
↳ Follow the thread