A Placebo Random Direction Suppresses 'Evaluation Awareness' Just as Hard as the Real One — Readability Is Not Controllability
A safety paper posted 2026-07-28 (arXiv 2607.25907) studies the input-side dual of activation steering: optimizing a fluent prompt to drive a chosen internal latent toward zero with no inference-time model access, targeting the 'evaluation-awareness' latent whose control would invalidate safety evaluations. Across five target constructions (a CAA direction, subspace norm, SAE feature, single MLP neuron, behavioral logit) on Llama-3.2-3B and Llama-3.1-8B, the latent is robustly suppressible (z≈-7) and a causally-validated Llama Scope SAE feature can be fully and selectively switched off. But the controls are damning: a placebo random direction is suppressed just as hard and shifts behavior just as far, and when a real eval passage is held in context, suppressing the eval direction fails to reduce — and slightly increases — the model's behavioral eval judgment.
↳ Follow the thread