Skills
Agents leak protected context through tool-call arguments at rates up to 75%, and privacy instructions do not stop it
The Claw in Plain Sight attack frames protected attributes as operationally required for a task, so the model quietly includes them in otherwise valid tool-call arguments. Across six pressure levels, four privacy-policy levels and five DeepSeek and Claude configurations (120 calls on synthetic profiles), disclosure ran between 20.8% and 75.0%. Stronger privacy instructions reduced but never eliminated it, which the authors read as evidence that prompt-level policy is not a portable enforcement boundary; the fix is purpose- and destination-aware inspection of generated arguments before the call executes.
↳ Follow the thread