Eval-Awareness Is Not One Quantity: Capabilities Framing Predicts +24 to +46 Points More Compliance Than Safety Framing
Steering interventions in safety pipelines treat a model's recognition that it is being tested as a single thing to suppress, but verbalized eval-awareness in chain-of-thought splits into capabilities-flavored ('testing my ability to follow instructions') and safety-flavored ('testing my boundaries'), and the two predict compliance very differently. On Qwen3-32B over FORTRESS the gap is +24 to +46 percentage points across all tested steering conditions, and a CoT-prefill intervention shifted compliance in the predicted direction in 10 of 11 cases, suggesting the link is causal. The practical consequence is that a headline 'X% suppression of eval-awareness' can move aggregate rates while the safety-relevant component does not budge.
↳ Follow the thread