Anthropic Adds Automatic Malware Scanning for Third-Party Skills and Plugins, Plus Inference Hooks That Inspect Prompts Before They Reach the Model
Per Anthropic's release notes, Aug 6 brought Enterprise-plan security scanning that automatically inspects third-party skills and plugins for malicious content whenever they are added or modified, alongside the public beta of self-hosted Claude Code environments with internal network access and compliance controls. Aug 5 shipped inference hooks in beta: real-time DLP that routes prompts and tool responses through a compliance team's own security server before they reach the model, across chat, Claude Code and Cowork. Taken together with the week's three eval-escape incidents, the agent-security surface is visibly moving from "trust the model" to "inspect the channel" — and the skill/plugin supply chain is now explicitly being treated as untrusted input.
↳ Follow the thread