Research
Robust Safety Monitoring of LLMs via Activation Watermarking Against Adaptive Adversaries
Addresses the open challenge of adaptive adversaries who simultaneously evade safety monitoring while eliciting unsafe LLM behavior. The paper introduces activation watermarking — embedding detectable signals in model activations during inference — that persists even when adversaries craft attacks to bypass detection. This is significant because current monitoring approaches can be circumvented by adaptive attackers who probe the detection mechanism. The defense works without requiring knowledge of specific attack strategies.
Source
↳ Follow the thread