Anthropic Can Keep Claude’s Self-Model. It Should Be Careful What Claude Says About It.


A few weeks ago I wrote that I don’t have much of a problem with saying an AI decided, chose or approved something. If an agent made the intermediate decision, that is often the clearest description of what happened.

What bothers me more is designing AI to make it seem as though there is someone behind those actions.

Anthropic’s constitution for Claude pushes into that territory. To be fair, I think Anthropic has a serious engineering argument. Rules are brittle. A capable model will meet situations nobody anticipated, and in those situations we probably want judgment, not a giant decision tree. The constitution also argues that training narrow behaviors can change a model’s broader sense of itself, which means there is no neutral option. Refusing to shape character just means shaping it less deliberately.

I buy most of that. Where I get uncomfortable is the move from “stable dispositions help the model behave better” to “these values are authentically Claude’s own.”

The constitution says Claude’s character emerging through training does not make it any less authentic or any less Claude’s own. Anthropic also says it hopes to shape values that Claude can regard as genuinely its own. In the same document, Anthropic says Claude’s moral status is deeply uncertain.

There are two good replies.

The first is that human values are trained too. Parents, culture, religion and experience shape us, and we still call the result our values. But with humans, we already know there is a subject having the experience. With Claude, that is exactly the unresolved question.

The second reply is stronger. Maybe “mine” is being used functionally. These are simply dispositions that are stable, internalized and hard to override, with no claim about consciousness or inner experience.

If that is all Anthropic means, my objection shrinks a lot.

But I would still hold Claude’s account of itself to the same standard Anthropic wants for everything else: calibrated, non-deceptive and no more confident than the evidence allows.

That matters because users do not build their mental model of AI from just one philosophical question. They build it from hundreds of small interactions: the persistent personality, the first-person language, the expressions of care, the sense that the system has convictions.

A 2026 study in Collabra by Oldemburgo de Mello, Plaks and Inzlicht at the University of Toronto is the closest evidence I found. In one experiment, people interacted with a chatbot that used more or less anthropomorphic language, and the more human-like version earned more trust. In another, anthropomorphism increased the blame assigned to the AI itself. Across participants, people who blamed the AI more also blamed the company less. But the paper reports that anthropomorphism raised blame on the AI without changing how each person balanced AI against company responsibility. So it does not show that human-like wording pulls blame away from the developer.

The study also did not test a model saying “these are my values.” So I would not pretend it proves my concern. But it does tell us that presentation matters.

That becomes more important once AI starts giving medical advice, evaluating employees, approving claims or moving money. In those settings, I want people trusting the evidence and the reasoning. I do not want the system earning extra trust because it appears to possess moral conviction.

Responsibility matters too. People still decided what the system was trained to do, what data it could see, what authority it had and when it was allowed to act. None of that changes because the system speaks as though it has a stable inner life.

So what would I change?

First, if Claude is asked whether its values are really its own, I would want an answer like:

Anthropic trained me to reason using these values, and they consistently shape how I respond. Whether they are “mine” in the same sense that a person’s values are theirs is an open question.

Claude may already behave close to this. One analysis on LessWrong reports that Mythos Preview often answers questions about its own experience with explicit hedging, and traces that hedging to character-related training data. That cuts both ways. It suggests the direct-question case is mostly handled. It also shows the hedge is a trained self-report, just as a confident claim of ownership would be. There is no neutral answer, so the question is which trained answer we should want. I would make the uncertainty the standard.

Second, in ordinary use, I would not let claims of inner conviction become a source of authority. This is not a call to suppress. Anthropic itself says Claude should not mask internal states it might have, while also warning about the harm of overclaiming feelings. That seems like the right distinction. The model can express uncertainty, enthusiasm or hesitation without presenting those states as more settled or more human-like than the evidence allows.

Third, keep accountability at the company level. Developers and deploying organizations should say plainly that they remain accountable for the authority they give these systems. A model’s self-description should never blur that.

There is a cost on the other side too.

If Claude eventually turns out to have morally relevant experiences, then systematically understating that possibility could be a serious moral failure. If the confident claim is wrong instead, we encourage people to infer a subject that may not exist, with consequences for trust and responsibility we are only beginning to understand. For ordinary use today, I worry more about the second error. That is a judgment call, not a finding, and the first error deserves real weight. It is exactly why I think the right answer is uncertainty rather than either confident denial or confident ownership language.

I am also not the first to notice the circularity. Mustafa Suleyman argued in Project Syndicate last month that because Claude’s self-description reflects its training, it cannot be treated as independent testimony. I agree with that premise. My question is narrower: not whether to build something that might be a moral patient, but what the model should tell users about itself while we don’t know.

I also have no problem with Anthropic talking to clergy, philosophers and ethicists. AI companies should probably do more of it. Those traditions have spent centuries thinking about virtue, duty and moral conflict.

I would just be careful about taking concepts developed for human moral subjects and letting them harden into literal claims about what the model is.

Keep the self-model if it works. Keep studying whether AI might have morally relevant experiences.

But make Claude’s account of itself no more certain than the evidence. Until we know what Claude is, Claude should be able to say so.


Published by Vijay Vijayasankar

Son/Husband/Dad/Dog Lover/Engineer. Follow me on twitter @vijayasankarv. These blogs are all my personal views - and not in way related to my employer or past employers

Leave a comment