Logical Validity Is Almost Perfectly Decodable From Hidden States Models Cannot Act On
Using matched valid-invalid premise-claim pairs varied across inference families, semantic domains, templates and difficulty, the authors probe five open-weight transformers and find that despite near-chance behavioral performance, logical validity is often almost perfectly decodable from hidden states, and stays strongly decodable under held-out templates, domains and inference families — including on examples the model answers incorrectly. But exhaustive leave-one-out tests reveal clear limits to that generalization, and interventions along probe-derived validity directions have only weak, nonspecific effects relative to random controls. The conclusion for anyone building probe-based monitors: representing a property, expressing it in behavior, and using it causally are three distinct things.
↳ Follow the thread