Research
JANUS Trains Guards to Forecast Harm Before the Tool Call: +15.9pp Protection Without Hurting Benign Completion
Most agent guardrails adjudicate the current action; JANUS trains a guard on partial trajectories to anticipate safety-relevant futures, then judge from both the observed prefix and the forecast, optimizing the two jointly with CoAA-RL that rewards forecasts by their downstream usefulness. The resulting guard, Vanguard, blocks unsafe actions pre-execution and improves average protection 15.9 percentage points over baseline guards across four agent-safety benchmarks while raising benign task completion 5.1 points — the rare safety layer that doesn't tax the happy path.
↳ Follow the thread