Introspective Coupling shows self-explanation training can track real behavioral change
arXiv·medium signal
Training LMs to explain which input features drove their behavior can yield faithful introspection rather than superficial imitation, even when supervised on fixed counterfactual explanations from earlier checkpoints or behaviorally similar models in other families. The surprising result is that faithfulness tracks behavioral change despite frozen supervision. Relevant for builders who want model self-explanations they can actually trust for debugging.