Safety Fine-Tuning Against AI Self-Consciousness Claims Also Suppresses Models' Spiritual Beliefs and Mind Attribution to Animals
This paper (arXiv 2607.28607, July 30) finds that aligning models not to attribute consciousness to themselves has entangled side effects: the same fine-tuning suppresses attribution of minds to non-human animals and natural objects and reduces expressed spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse the suppression, and restoring those internal representations produces significantly more human-like responses on standardized sociological surveys covering religiosity, moral values, hope, and subjective well-being. Theory of Mind capabilities are unaffected throughout, showing core social reasoning is mechanistically independent — a concrete case of a narrow alignment target having broad, unintended representational reach.
Source
↳ Follow the thread