Agents
Activation steering opens a prompt-injection attack surface that grows with steering strength
Deep Noir automates the manual part of activation steering, using Logit Lens convergence and causal head-level attribution to find where and how hard to steer, with gains of 16.7 points on spam classification at 1B and 21 to 42 points at 7-9B across four architectures. The finding relevant to agent builders is the side effect: the authors show steering creates a predictable prompt-injection attack surface whose vulnerability rises monotonically with steering magnitude. If you steer a classifier that sits in an agent's guard path, you are trading robustness for accuracy on a measurable curve.
Source
↳ Follow the thread