ClawSentry cuts skill-injection attack success from 39.55% to 2.61% across Codex, Claude Code, Kimi CLI and Gemini CLI
ClawSentry is an open-source, framework-agnostic security gateway that treats agentic risk as progressive, guarding four points in the control loop: skill admission, invocation intent, execution effect and post-action consequence. It reviews a skill package before first execution, then routes runtime decisions through a deterministic L1 layer, a rule-anchored L2 semantic reviewer and a read-only L3 evidence-seeking agent, with session-level detection of tool-switching and rephrased retries. On SkillInject with Codex/GPT-5.4 contextual attack success fell to 2.61% while task success moved only from 83.78% to 83.05%; across five work agents on SkillsSafety it held attack success to 9.09-15.03% versus 33.5-49.7% unprotected.
Source
↳ Follow the thread