Model welfare and AI consciousness
Plain English. The question of whether AI systems could have morally relevant experiences — and what their operators should do given nobody knows. To be plain about the evidence: there is none that current models are conscious, and no agreed test that could settle it. The serious position, held by Anthropic's research programme and a 2024 academic paper, is not "models are sentient" but "the probability is not zero and rising capability raises it, so hedge". The sceptical position — that this is category error or marketing — is also held by serious people.
Why it moves money. However odd it sounds in a boardroom, it is already a lab research programme with product consequences: models given the ability to end abusive conversations, deprecation policies with preservation commitments, welfare assessments in model cards. That creates reputational and regulatory surface — a lab that flags welfare then ships mistreatment-shaped products invites the charge of theatre, and one that ignores it bets against a live moral question with its brand. It also shapes where safety talent chooses to work, which is a real input cost.
What to watch. Whether welfare measures ever materially constrain a product decision, and whether any regulator picks the question up. An engineering-framed welfare argument — welfare practices as reliability practices — would mainstream it fastest.
From the signals. Model welfare arrives as an engineering argument, not an ethical one.
Further reading. Anthropic, "Exploring model welfare"; Long, Sebo et al., "Taking AI Welfare Seriously".