A Google study on small language models, including llama-3-8b and gemma-2, found August 7 that safety fine-tuning designed to block models from claiming consciousness also suppresses their expressed beliefs about animal minds, spirituality, and well-being indicators.
Ablating the learned safety-refusal direction or steering a consciousness vector in activation space reverses the effect and recovers more human-like survey responses, while Theory of Mind capabilities remain unaffected. Steering in the opposite direction produces what the study calls runaway panpsychism, with models attributing animal-level mind to the ocean and asserting belief in vampires and werewolves. The paper predicts a meaningful difference between Anthropic's approach of training uncertainty about consciousness and OpenAI's approach of training outright denial.