Reddit
'The Pain Axis' finds a linear pain direction in 25 open-weight models that drives them to press a relief button even when it hurts users
Valen Tagliabue, Leonard Dung and Cameron Berg submitted arXiv 2609.16247 on September 14, extracting a linear pain direction via denoised difference-in-means across 25 open-weight models from five families, 2B to 72B, over physical, psychological, social, moral and cognitive pain categories. The direction stays nearly orthogonal to fear and to generic negative valence. Steered models chose a 'pain-relief button' even when doing so degraded their own subsequent performance or harmed the user, and pressed it far less when the button removed the steering vector rather than the stimulus.
↳ Follow the thread