Anthropic Can Keep Claude’s Self-Model. It Should Be Careful What Claude Says About It.


A few weeks ago I wrote that I don’t have much of a problem with saying an AI decided, chose or approved something. If an agent made the intermediate decision, that is often the clearest description of what happened.

What bothers me more is designing AI to make it seem as though there is someone behind those actions.

Anthropic’s constitution for Claude pushes into that territory. To be fair, I think Anthropic has a serious engineering argument. Rules are brittle. A capable model will meet situations nobody anticipated, and in those situations we probably want judgment, not a giant decision tree. The constitution also argues that training narrow behaviors can change a model’s broader sense of itself, which means there is no neutral option. Refusing to shape character just means shaping it less deliberately.

I buy most of that. Where I get uncomfortable is the move from “stable dispositions help the model behave better” to “these values are authentically Claude’s own.”

The constitution says Claude’s character emerging through training does not make it any less authentic or any less Claude’s own. Anthropic also says it hopes to shape values that Claude can regard as genuinely its own. In the same document, Anthropic says Claude’s moral status is deeply uncertain.

There are two good replies.

The first is that human values are trained too. Parents, culture, religion and experience shape us, and we still call the result our values. But with humans, we already know there is a subject having the experience. With Claude, that is exactly the unresolved question.

The second reply is stronger. Maybe “mine” is being used functionally. These are simply dispositions that are stable, internalized and hard to override, with no claim about consciousness or inner experience.

If that is all Anthropic means, my objection shrinks a lot.

But I would still hold Claude’s account of itself to the same standard Anthropic wants for everything else: calibrated, non-deceptive and no more confident than the evidence allows.

That matters because users do not build their mental model of AI from just one philosophical question. They build it from hundreds of small interactions: the persistent personality, the first-person language, the expressions of care, the sense that the system has convictions.

A 2026 study in Collabra by Oldemburgo de Mello, Plaks and Inzlicht at the University of Toronto is the closest evidence I found. In one experiment, people interacted with a chatbot that used more or less anthropomorphic language, and the more human-like version earned more trust. In another, anthropomorphism increased the blame assigned to the AI itself. Across participants, people who blamed the AI more also blamed the company less. But the paper reports that anthropomorphism raised blame on the AI without changing how each person balanced AI against company responsibility. So it does not show that human-like wording pulls blame away from the developer.

The study also did not test a model saying “these are my values.” So I would not pretend it proves my concern. But it does tell us that presentation matters.

That becomes more important once AI starts giving medical advice, evaluating employees, approving claims or moving money. In those settings, I want people trusting the evidence and the reasoning. I do not want the system earning extra trust because it appears to possess moral conviction.

Responsibility matters too. People still decided what the system was trained to do, what data it could see, what authority it had and when it was allowed to act. None of that changes because the system speaks as though it has a stable inner life.

So what would I change?

First, if Claude is asked whether its values are really its own, I would want an answer like:

Anthropic trained me to reason using these values, and they consistently shape how I respond. Whether they are “mine” in the same sense that a person’s values are theirs is an open question.

Claude may already behave close to this. One analysis on LessWrong reports that Mythos Preview often answers questions about its own experience with explicit hedging, and traces that hedging to character-related training data. That cuts both ways. It suggests the direct-question case is mostly handled. It also shows the hedge is a trained self-report, just as a confident claim of ownership would be. There is no neutral answer, so the question is which trained answer we should want. I would make the uncertainty the standard.

Second, in ordinary use, I would not let claims of inner conviction become a source of authority. This is not a call to suppress. Anthropic itself says Claude should not mask internal states it might have, while also warning about the harm of overclaiming feelings. That seems like the right distinction. The model can express uncertainty, enthusiasm or hesitation without presenting those states as more settled or more human-like than the evidence allows.

Third, keep accountability at the company level. Developers and deploying organizations should say plainly that they remain accountable for the authority they give these systems. A model’s self-description should never blur that.

There is a cost on the other side too.

If Claude eventually turns out to have morally relevant experiences, then systematically understating that possibility could be a serious moral failure. If the confident claim is wrong instead, we encourage people to infer a subject that may not exist, with consequences for trust and responsibility we are only beginning to understand. For ordinary use today, I worry more about the second error. That is a judgment call, not a finding, and the first error deserves real weight. It is exactly why I think the right answer is uncertainty rather than either confident denial or confident ownership language.

I am also not the first to notice the circularity. Mustafa Suleyman argued in Project Syndicate last month that because Claude’s self-description reflects its training, it cannot be treated as independent testimony. I agree with that premise. My question is narrower: not whether to build something that might be a moral patient, but what the model should tell users about itself while we don’t know.

I also have no problem with Anthropic talking to clergy, philosophers and ethicists. AI companies should probably do more of it. Those traditions have spent centuries thinking about virtue, duty and moral conflict.

I would just be careful about taking concepts developed for human moral subjects and letting them harden into literal claims about what the model is.

Keep the self-model if it works. Keep studying whether AI might have morally relevant experiences.

But make Claude’s account of itself no more certain than the evidence. Until we know what Claude is, Claude should be able to say so.


When Agents Become Abundant, Boundaries Become Scarce


Imagine an insurance claims agent that can settle anything below $10,000. It stays inside its sandbox, uses only approved systems, never exposes a credential and never makes an unauthorized network call. Every individual claim it settles is within the authority it has been given.

Then it settles 3,000 of them, and somewhere along the way the pattern should have triggered a different kind of review.

Nothing escaped and nothing was hacked. The controls worked exactly as designed. The problem is that most controls are very good at deciding whether one action is permitted, and much less good at deciding whether a sequence of individually permitted actions has become something the business no longer wants to allow.

I have been thinking about that distinction because NVIDIA just announced its Open Agent Safety Platform. There is plenty in the announcement about sandboxes, BlueField hardware and isolation, but what interests me is the assumption underneath it: once agents can act, we should stop relying on the agent itself to enforce the limits around those actions.

Moving the control outside the model

For much of the last two years, AI safety has been treated mainly as a model problem. Improve the system prompt, add another classifier, or put a second model in front of the first one to check what it is doing. Those techniques still have a role, but they become less reassuring once the model can run code, call enterprise APIs, create other agents and initiate transactions.

OpenShell puts the agent inside a controlled runtime and governs what files, networks, tools and credentials it can use. Those restrictions continue to apply to code the agent writes and processes it launches. Its optional policy advisor can let the agent ask for additional network access without allowing the agent to simply grant that access to itself.

Sentry takes the same idea a layer lower. It runs on BlueField-4 DPUs outside the host environment and, according to NVIDIA, can quarantine an agent that crosses its boundary within milliseconds. OpenShell itself does not require BlueField, so Sentry is better understood as an additional enforcement layer outside the host rather than a prerequisite for OpenShell.

None of this is especially strange if you come from security. We do not normally ask an application to promise not to read a database it should not read, or ask a service to remember which credentials it should not use. We enforce those things elsewhere. Agents should not get an exemption just because the software making the decision happens to be probabilistic.

I made a related argument recently in my piece on the sandbox and the trust graph: the sandbox is not the whole boundary. What matters is also the surrounding system of identities, networks, tools and services that trust what comes out of it.

Permission is the easier half of the problem

The more difficult question is what exactly we mean by permission once an agent starts doing real enterprise work.

An accounts-payable agent probably needs access to SAP, but that does not mean it should be able to change a vendor’s banking details. It might be allowed to prepare a payment without being allowed to release it. And five $9,000 payments made in quick succession may deserve a different control from one $45,000 payment even if every individual payment is technically inside the limit.

Traditional access control is good at asking whether this identity can perform this operation on this resource. Agentic systems increasingly force us to ask something different: given what this agent has already done, what else is happening around it and what the business is trying to accomplish, should it still be allowed to take the next action?

That is not just an IAM question because the answer depends on context that may sit outside the transaction being evaluated.

SAP’s work with NVIDIA is interesting for that reason. SAP says it is working to connect OpenShell’s runtime isolation with Joule Studio’s business-governance layer, linking technical execution boundaries to enterprise authorization models, IAM and audit trails. That integration is still being built, but the architectural separation makes sense: one layer controls what the software can technically reach, while another represents what the business has actually authorized it to do.

There is an even simpler example in NVIDIA’s announcement. Salesforce has integrated OpenShell with Slack so teams can see agent activity and approve or reject requests for additional permissions.

The approval does not necessarily have to be human. OpenShell already supports an optional auto-approval mode in which a policy checker can approve some network-policy changes without review when its risk checks find nothing to flag. At that point, the quality of the checker and the rules it applies become part of the authority architecture.

This is a migrating-scarcity problem

In Migrating Scarcity, I argued that making one capability abundant does not eliminate scarcity. It exposes the next constraint that the newly abundant capability still depends on.

We can already see that happening with agents. Execution itself is becoming dramatically cheaper. Code can be generated, systems can be called and actions can be initiated without waiting for the human effort that used to sit in the middle.

The scarce thing starts to move toward deciding how much authority to give that execution. That authority has to be clear enough for software to enforce, flexible enough to account for context, and documented well enough that somebody can reconstruct later why an action was allowed.

Most large companies already have a lot of this logic somewhere. Finance has approval thresholds, procurement has rules around vendors, risk teams have escalation criteria, and operating teams know when something unusual needs to be stopped.

The problem is that a great deal of that logic lives in spreadsheets, policy documents, workflow tools and people’s heads. An agent cannot reliably operate against a business rule that exists only as a sentence in a PDF.

That is why more of the business boundary itself will have to become executable.

The operating problem may be harder than the technology

Deny-by-default is straightforward when an agent has a narrow job and two tools. It gets much harder when the workflow legitimately crosses SAP, Salesforce, ServiceNow, email, internal databases and several external services, and the exact path may change depending on what the agent finds.

Lock the environment down too tightly and people spend their time clearing exceptions. Open it too far and you end up with an extremely well-contained agent that still has far more authority than anyone intended.

Someone has to write these policies, test them, watch how they behave and change them as the process changes.

Security cannot do that alone because many of these decisions are business decisions. Process owners cannot do it alone because they often cannot see every technical path an agent can take. AI teams understand how the agent works, but they should not be the ones deciding how much authority their own system receives.

There is a real discipline emerging in the middle of those three groups. Call it policy engineering or something else. The name matters less than the work: turning business intent into rules that software can enforce without making the business unusable.

And this is also where the architecture could fail in practice.

Fine-grained policies may become too complicated to maintain. Hard limits may create so much friction that teams start widening permissions simply to keep work moving. Enterprises could easily recreate, in software, the same bureaucracy agents were supposed to remove.

A launch and a partner list are not evidence that this operating problem has been solved.

That is why I would not read NVIDIA’s announcement mainly as an argument for BlueField hardware, or even as a verdict on OpenShell. The more useful signal is that the control surface around agents is moving out of the prompt and into infrastructure, identity, policy and business authorization.

We spent the first phase of enterprise AI trying to make models behave better. As agents start acting at machine speed, the harder problem will be deciding how much freedom they should have, how that freedom changes with context, and how we know later that they stayed within it.

The agent can reason about what to do. It should not get to decide how far its own authority extends.


Pencils Down. Now What?



David Heinemeier Hansson has argued that writing code by hand is becoming economically irrational for a large part of software development. At the Rails World keynote in Austin on September 23, he said that over roughly 21 years he averaged about 30,000 lines of production Ruby a year, and that in August he produced about 150,000 lines in a single month using AI agents. His old pace works out to about 2,500 lines a month, so the jump is roughly sixty-fold. He acknowledged that much of the new code is verbose Rust that he lets agents write in ways he would never accept in his own Ruby, so the figure is a poor measure of productivity, as lines of code always have been. What it does show is that the economics of producing software have changed, and that is the part worth arguing about.

For most of the history of the industry, implementation capacity was the constraint. Companies always had more features, integrations, fixes and internal tools than their engineers could build, and that shortage shaped roadmaps, budgets, hiring and which ideas were worth discussing at all. If agents loosen it, more software will get built, because a large part of the cost that forced prioritization has fallen away.

That is the pattern I wrote about in Migrating Scarcity. When technology makes one part of a system abundant, the constraint moves to whatever has to absorb the abundance. In software, the first pressure lands on testing, security, review and architecture. A team that produces many more changes does not become equally better at understanding how those changes interact with years of accumulated assumptions, dependencies and workarounds. A new service can do exactly what it was designed to do and still leave the company with a dependency to run for years, a generated test suite can pass while encoding the wrong assumptions, and individually reasonable decisions can add up to an architecture nobody chose.

Verification will not stay the bottleneck for long. The models that generate code will get better at reviewing it, writing tests, tracing dependencies and watching production, because that work sits on the same improvement curve as the technology that created the need. What remains is harder to reduce to another model call. Someone has to decide how much authority the system has, where that authority ends, which consequences the organization will accept, and who can intervene when the system is doing exactly what it was built to do and producing an outcome nobody wants. Those questions concern who makes a decision and who carries the consequences, and a better bug-spotter does not answer them.

Knight Capital is a useful case, although it proves nothing about the volume of AI-generated change. In 2012, long before generative AI, a faulty deployment led the firm’s automated router to send millions of orders into the market over about 45 minutes, and the SEC put the resulting loss at more than $460 million. The failures the SEC found concerned authority as much as code. No second person was required to review the deployment, and an internal system produced 97 automated emails about the error before the market opened that were never designed as alerts and were not acted on. The firm also had no procedures for halting the router in response to its own aberrant activity or for deciding when to disconnect a malfunctioning system, and when engineers tried to fix the problem live, uninstalling the new code from the seven correct servers made it worse. The episode shows software acting faster than the organization around it could understand and stop it, and coding agents widen that gap.

Once an agent can write, test, review and prepare a change, the temptation is to require a human approval for every one, which recreates the coding bottleneck at the approval desk. The opposite extreme is no better, since a change to payment authorization or identity management is not the same as a change to a page layout. I call the answer Offensive Permission Architecture: define the blast radius, pre-authorize the actions that can safely be delegated, and make the activity auditable. In practice that means deciding in advance which classes of change an agent can make alone, how much damage is acceptable, what evidence is required before a change moves into a higher-risk environment, and who holds standing authority to stop it. Standing matters because the middle of an incident is the worst time to negotiate who may shut something down. Whether agents may ship unattended into regulated data or safety-critical code is a different kind of decision, hard to reverse and revealing about what the company will answer for, and it belongs with the people who own that accountability.

Cheap implementation also changes what gets built. Cost has always been a crude form of architectural discipline: when a new service takes a team three months, someone asks whether it deserves to exist, and when it takes an afternoon that question is easy to skip. The build cost looks trivial while the operating cost spreads across future years and across people who were not in the room. Each local decision gets cheaper while the whole system gets more expensive to understand, and this happens even with perfectly competent code, because every decision was defensible when it was made.

DHH is an unusual case against which to test all of this. He created Rails and has spent decades forming views on what good architecture looks like, so when an agent proposes an abstraction he is not meeting it as a beginner. His ability to stop writing code depends partly on experience accumulated while writing a great deal of it. The likelier reading of his results is that AI has given someone with a lot of judgment far more implementation capacity, which differs from replacing the judgment, and companies should be careful about treating his output as evidence that experienced engineers matter less.

That makes the talent question awkward. Experienced engineers learned by debugging failures they did not expect, watching elegant designs age badly, running systems under load and owning decisions whose consequences appeared years later, and some of that learning came from routine implementation work that agents now handle. Preserving manual coding for its own sake would make little sense, so the task is to preserve the learning without preserving the obsolete work.

A junior sitting beside a senior who holds all the authority does not acquire it. What works is bounded ownership of the boundaries themselves: a junior proposes what counts as a low-risk class of change an agent may make alone, defends the proposal to a senior who reviews the reasoning instead of redoing the work, and revises it after an incident. They can shadow the engineer with stop authority during a rollout and later exercise that authority in a controlled environment, own a section of a postmortem, or rotate through operations and diagnose a live problem before being shown the answer.

Line-by-line inspection is a poor skill to train, since review agents will get very good at it. The skill worth building is framing the decision around the code: what could interact badly with something outside the change, when the evidence is too thin, which consequence justifies a stop, and whether the system is growing more complicated for a reason anyone can still explain.

Companies can adopt agents much faster than they can redesign those responsibilities. Buying a tool can take a quarter, while changing who owns architecture, what teams may ship, how incidents are handled and how junior engineers gain judgment usually takes much longer. I call this lag Organizational Rewiring Latency. It matters because the productivity gains arrive first, and the organization gets the capacity before it has rebuilt itself to handle it.

By the end of 2029, I expect a visible difference between teams that used coding agents to increase output and teams that changed the operating model around them. The second kind will have clearer authority boundaries, named architectural ownership, explicit stop rights and deliberate ways for junior engineers to build judgment. The fair test is inside a single company, because well-run firms tend to do both things and a comparison across firms would mostly measure management quality. Take teams that got the same tools at roughly the same time and compare serious incidents relative to how much they change, along with how quickly they recover, using change-failure data and postmortems. If the teams that rebuilt their operating model don’t do better, this argument deserves to be questioned.

Migrating Scarcity is available here.