Fetching from the wire…
Public story · 2026-08-24 · high
ClawSentry checks skill admission, intent, execution and outcome, and task completion barely moved in testing.
Why now: The paper's arXiv listing places it in August 2026, with no word yet on whether any agent framework has adopted the gateway.
ClawSentry, an open-source framework-agnostic security gateway, cut the success rate of skill-injection attacks against AI agents from 39.55% to 2.61%, according to the paper introducing it. Teams deploying agents that touch real skills need proof a security layer doesn't cost them throughput. Task success held at 83.05% with the gateway running, just under the 83.78% baseline without it, on a benchmark called SkillsSafety.
It sits between an agent and the skills it invokes, checking four points in that loop: skill admission, invocation intent, execution effect, and post-action consequence. A deterministic first layer handles routine cases. A rule-anchored reviewer checks flagged actions against fixed criteria, and a third, read-only agent gathers evidence before ruling on anything left over.
Across five separate work agents tested on SkillsSafety, ClawSentry held attack success between 9.09% and 15.03%. Without it, the same agents failed between 33.5% and 49.7% of the time. The paper doesn't say whether any agent framework has integrated ClawSentry yet.
Each link below shares sources, entities, or timing with this story.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, attack, control, execution); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, effect, execution); pushes against this story (against).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, cost, success); traces where this leads (downstream).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (against, attack, success); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, attack); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, cost); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, control); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, agent, attack); pushes against this story (but).