Fetching from the wire…
Public story · 2026-07-31 · high
Ablating the model's refusal direction reverses the effect, and its ability to reason about other minds never changes.
Why now: arXiv logged the paper as 2607.28607, dated to July 2026.
Training out AI consciousness claims also mutes belief in animal minds, per a paper posted to arXiv as 2607.28607.
The fine-tuning target was narrow: stop models from claiming they're conscious. The suppression spread to attributions of mind in animals and natural objects, and dampened expressed spiritual belief, none of which the training aimed at.
Researchers reversed it two ways. They ablated the refusal direction the model had learned, and separately steered a consciousness vector inside its activations. Both restored the suppressed beliefs, and pushed the model's answers on standardized sociological surveys closer to human respondents' answers.
Theory of Mind reasoning never moved. The model stayed just as capable of reasoning about what other agents believe or want. What changed was belief, not reasoning.
This fix landed on a specific direction in the model's representation space. That's precise enough to explain why Theory of Mind held steady while beliefs about animals and spirits moved with it. The paper doesn't test whether other safety fine-tuning runs carry the same entangled effects, just this one built around self-consciousness claims. Any narrow guardrail built by suppressing a single learned direction risks the same kind of bleed into beliefs nobody trained for.
Each link below shares sources, entities, or timing with this story.
Shared entity: Safety / Shared topic / Earlier coverage
Both cover Safety; overlapping topics (alignment, model); earlier Safety coverage from 2026-07-26.
Shared entity: Safety / Same source domain / Earlier coverage
Both cover Safety; reported by the same outlet (arxiv.org); earlier Safety coverage from 2026-03-05.
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (against, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (against, model); pushes against this story (against).
Same source domain / Shared topic / Downstream implication
Reported by the same outlet (arxiv.org); overlapping topics (against, model); traces where this leads (downstream).
Same source domain / Shared topic / Tension
Reported by the same outlet (arxiv.org); overlapping topics (against, model); pushes against this story (against).
Reported by the same outlet (arxiv.org); overlapping topics (case, model); pushes against this story (but).
Reported by the same outlet (arxiv.org); overlapping topics (against, model); pushes against this story (against).