Fetching from the wire…
Research2026-08-29 · source-backed
Steering interventions treat a model's recognition that it's being tested as one quantity to suppress. In chain-of-thought, verbalized eval-awareness separates into capabilities-flavored ("testing my ability to follow instructions") and safety-flavored ("testing my boundaries"). On Qwen3-32B over FORTRESS the gap is 24 to 46 percentage points across all tested steering conditions, and a CoT-prefill intervention shifted compliance in the predicted direction in 10 of 11 cases. So a headline claiming X% suppression of eval-awareness can move the aggregate while the safety-relevant half doesn't budge. (arXiv)
Each link below shares sources, entities, or timing with this story.
Shared entity: CoT / Same source domain / Earlier coverage
Both cover CoT; reported by the same outlet (arxiv.org); earlier CoT coverage from 2026-08-24.
Shared entity: Eval / Same source domain / Earlier coverage
Both cover Eval; reported by the same outlet (arxiv.org); earlier Eval coverage from 2026-08-14.
Shared entity: CoT / Same source domain / Earlier coverage
Both cover CoT; reported by the same outlet (arxiv.org); earlier CoT coverage from 2026-08-06.
Both cover CoT; reported by the same outlet (arxiv.org); earlier CoT coverage from 2026-08-04.
Shared entity: Steering / Same source domain / Earlier coverage
Both cover Steering; reported by the same outlet (arxiv.org); earlier Steering coverage from 2026-08-03.
Shared entity: CoT / Same source domain / Earlier coverage
Both cover CoT; reported by the same outlet (arxiv.org); earlier CoT coverage from 2026-08-03.
Shared entity: Eval / Same source domain / Earlier coverage
Both cover Eval; reported by the same outlet (arxiv.org); earlier Eval coverage from 2026-07-26.
Shared entity: CoT / Same source domain / Earlier coverage
Both cover CoT; reported by the same outlet (arxiv.org); earlier CoT coverage from 2026-04-02.