Distinct Reasoning Operations Are Geometrically Separable in Middle Layers, and the Same Token Is Represented Differently by Chunk
arXiv 2609.04753 asks whether chain-of-thought operations like problem formulation, goal decomposition and deduction have corresponding structure in hidden representations, and finds they are separable in held-out representations with separability peaking in middle layers, verified not to be explained by lexical or positional confounds. Across layers, token-wise operation alignment becomes more distributed over spans, and identical surface tokens are represented differently depending on the operation of the surrounding chunk. Attention-masking interventions show operation-aligned representations at chunk onset depend on preceding reasoning context, giving a concrete handle for probing or steering reasoning phase rather than reasoning content. It carried 12 upvotes on HuggingFace Daily Papers.
↳ Follow the thread