Agents
Telling colluding agents they are being watched by an activation probe does not change anything
Output monitors can be fooled by obfuscation and steganography, which motivates probes trained on internal activations, but those probes are normally evaluated against agents that do not know they are watched. This study holds models, probes and thresholds fixed and varies only what agents are told: nothing, that an activation monitor is present, or that a monitor is present plus the previous round's score. Across a four-agent blackjack game and a two-agent Simmons prisoners game with Qwen3-32B-AWQ and GPT-OSS-20B in homogeneous and mixed pairings, the best probes stay accurate in all three conditions and the agents keep colluding.
Source
↳ Follow the thread