Research
MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning in CoT Models
MirageBackdoor demonstrates a novel backdoor attack on Chain-of-Thought reasoning where the model produces correct-looking reasoning chains but arrives at wrong final answers. Unlike prior CoT backdoors that corrupt the reasoning itself, this attack preserves coherent intermediate steps while manipulating only the conclusion — making detection significantly harder. Directly relevant to anyone deploying CoT-based agents in production.
Source
↳ Follow the thread