LLMs Make Systematic Errors on Simple Counterfactuals About Their Own Behavior, and RL Fixes the Score Without Granting Introspection
This EMNLP 2026 paper benchmarks self-modeling using verifiable behavioral questions, such as whether a given prompt edit would change the model's own final answer, and finds current models show non-trivial but limited skill with systematic failures on simple counterfactuals about themselves. A scalable synthetic-data pipeline plus reinforcement learning raises aggregate self-modeling across three open-source model families with some transfer to held-out tasks. The authors are careful about what that buys: improved self-modeling does not consistently amount to introspection, and may not come from privileged access to the model's internal decision process, which matters for anyone relying on a model's self-reported confidence or self-critique.
↳ Follow the thread