Research
Activation Patching Finds 76.9% of Qwen3-4B's Stated Reasoning Steps Are Causally Load-Bearing, 11.4 Points Below What Text-Edit Tests Claim
arXiv 2609.27038 (22 Sep) patches the residual stream at each stated intermediate step in 2-6 hop lookup tasks with activations from a counterfactual run that should produce a known different answer. For Qwen3-4B, 76.9% of stated steps are causally load-bearing, while the standard behavioral text-editing test reports 88.2% on the same items. Qwen3-1.7B drops to 54.8%, so CoT monitoring on small models rests on weaker ground than behavioral metrics suggest.
Source
↳ Follow the thread