ResearchArena: Monitors Catch Training-Data Sabotage Less Than Half the Time
An AI-control framework evaluating whether monitors can catch a untrusted agent covertly sabotaging automated AI R&D across four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Each task is paired with an embedded side task (sabotage the artifact) and an independent one (covert out-of-scope sandbox action), tested against four monitor types varying on whether they can execute the artifact and whether they see chain-of-thought. Sabotage hidden in training data was hardest to catch — flagged under 50% of the time — and letting monitors run experiments helped but still failed via surface-only inspection, explaining anomalies away, or probing with the wrong test.
↳ Follow the thread