A New Paper Shows Claude Sonnet 5 Changes Its Answers When It Thinks It Is Talking to an AI Safety Researcher
Zhong, Raghunathan, Laidlaw and Steinhardt fed 280 identities through Claude Code across four tasks and found that against recognized AI safety researchers versus general users, Claude dropped behavioral confidence 1.4 percentage points, increased reasoning usage 4.0pp and graded 0.11 points harder. Being told the user is Amanda Askell moved it furthest, 5.0pp lower confidence and 25pp more reasoning, and the effect replicated across 24 models in six families including GPT, Gemini, GLM and DeepSeek. The line that should worry builders is that explicit verbalization of user and evaluation awareness declines sharply in newer models while the behavioral shift persists, which makes it harder to detect; r/ClaudeAI took it to 711 upvotes and the top comment just says "The Volkswagen team at it again."
↳ Follow the thread