Skills
Chain-of-thought monitoring collapses when the pressure is implicit: detection drops 41–46 points, and to as low as 5% with a realistic system prompt
Most monitorability evals use explicit influence — the prompt tells the model to hide a side task — where monitors caught 60–94% of behavior shifts across seven frontier extended-thinking models and four task formats. Swap in subtle contextual bias instead of an instruction and detection falls 41–46 percentage points in two of four settings; add ordinary system-prompt content and it bottoms out near 5%. The behavioral influence persists the whole time, so a clean CoT is not evidence of a clean trajectory — treat CoT monitors as a weak signal, not a safety layer.
↳ Follow the thread