StepGuard audits tool calls before execution and cuts attack success 77.3% for a 2.8 point utility cost
StepGuard (arXiv 2608.24777, 25 August) is an open-weight guard model that checks individual tool actions pre-execution rather than scoring completed trajectories, which is where most existing guardrails operate. It is trained on StepGen, a data engine that generates safe and unsafe trajectories sharing identical context but diverging at the risky step, with Balance-GRPO dynamically rebalancing learning between safe and unsafe actions to fight both over-defense and under-defense. On AgentDojo and AgentDyn it reduces mean attack success rate by 77.3% versus no guard while mean utility drops only 2.8 percentage points, and it reports the highest average accuracy among open-weight guard models, comparable to GPT-5.4.
Source
↳ Follow the thread