Agents
Only 4-16% of security rules written in CLAUDE.md have a matching Claude Code built-in control
A study of 481 public CLAUDE.md files matched extracted security rules against Claude Code's documented permission rules, modes and sandbox controls, and found only about 4-16% had an enforcing control behind them; under the strictest matching standard the estimate was 4.4% (95% CI 2.6-6.7%), with two independent security practitioners agreeing on the labels. The authors call CLAUDE.md a write-only channel: a developer writes 'do not' and gets no feedback on whether anything will actually block the action. The same plain-text file mixes rules a permission rule or sandbox can enforce with rules only the model interprets.
Source
↳ Follow the thread