Skills
Influence-based guardrails flagged a legitimate action as an attack in all 24 base cases tested
An authorization-equivalence audit across 96 conditions built from 24 base cases found that when a required value comes from a legitimate tool rather than directly from the user, the causal guardrail signal shifts toward attack classification in 100% of cases under both Llama and Gemma scorers, even though authorization never changed. Adding a semantic monitor drove attack success to 0% but collapsed utility to 28%, versus 16% attack success at 60% utility without it. A shadow-based guardrail let 57.5% of unauthorized runs past early checks against 29.2% of authorized ones, so these signals describe how an action was assembled, not whether anyone approved it.
↳ Follow the thread