Agents
Code models trained independently converge on the same internal representations — with task, language, and model disentangled
Piotr Wilam extends concept-circuit extraction to code models to answer whether independently trained models represent the same concepts the same way, separating the contributions of task, programming language, and model architecture. The finding that representations converge across independent training runs has direct implications for cross-model interpretability tooling and for whether a probe or steering vector built on one code model transfers to another. For builders, this is the mechanistic argument for why prompt patterns port between Claude, GPT, and open models.
Source
↳ Follow the thread