Research
ThoughtSteer: Backdoor Attacks on Continuous Latent Reasoning Achieve ≥99% Success Rate, Evade All Defenses
Researchers demonstrate ThoughtSteer, an attack that perturbs a single embedding vector in models that reason in continuous hidden states (Coconut, SimCoT), achieving ≥99% attack success with near-baseline clean accuracy across 124M–3B parameter scales. The attack transfers to held-out benchmarks without retraining (94–100%), survives 25 epochs of clean fine-tuning, and evades all five evaluated active defenses. Neural Collapse in the latent space creates a geometric attractor that amplifies the perturbation — individual vectors still encode the correct answer, but the collective trajectory produces the wrong one.
Source
↳ Follow the thread