Dispatch
James Mickens argues chain-of-thought monitoring can never be a sound security control
A paper submitted 2 September (arXiv 2609.02852) introduces 'linguistic illegibility,' the gap between what a model says and the math it actually performs over activation spaces. Mickens argues that chain-of-thought monitoring, constitutional self-critique and activation probing are unsound in principle as security mechanisms, and that sandboxing must rest on techniques independent of what the model reports, specifically taint tracking over which system states a model's output influenced. For builders, the practical read is to spend effort on the sandbox boundary rather than on reading agent reasoning traces for safety.
Source
↳ Follow the thread