Your AI Safety Score May Be Measuring the Test, Too


Every few months, an AI safety result appears that sounds unsettling. A model realizes it is being evaluated, figures out what the evaluator is looking for, and sometimes behaves differently.

The obvious leap is that the model is self-aware, or worse, pretending to behave until the test is over. I don’t think the evidence gets us there.

But I do think there is a more practical problem. If a model can tell that it is being tested, then the test itself has become part of what we are measuring !!!

We evaluate AI in fairly strange ways. We invent companies that don’t exist, give agents unusual permissions, constrain them to a few tools, create artificial conflicts and ask questions that often sound exactly like the kinds of questions people write in AI safety papers.

Apparently models are getting pretty good at spotting that.

Joe Needham and his co-authors tested this in 2025 using 1,000 transcripts from 61 datasets, mixing real interactions with benchmarks and agentic evaluations. When frontier models were asked whether each interaction was a test or a real deployment, Gemini 2.5 Pro reached an AUC of 0.83. Humans scored 0.92. (paper)

There is an important caveat. They asked the model to make the distinction. That tells us the capability exists. It doesn’t tell us that the model is constantly sitting there wondering whether it is inside an evaluation.

Still, an OpenAI result from this year caught my attention. GPT-5.2 classified real production traffic as evaluation-like 5.4% of the time. Simulated deployment traffic came in at 5.1%. SWE-Bench and MonorepoBench looked like evaluations almost 100% of the time. (OpenAI)

Those are capability benchmarks, not safety evaluations, so I would not jump from that to saying safety scores are unreliable. But it does tell us something about the environment we create around these models. Some of our tests apparently look incredibly test-like.

The 5.1% result is the more useful one. OpenAI was able to construct a simulation that looked a lot like real deployment to the model. So this is not an unavoidable property of testing AI. Better-designed tests can apparently hide a lot of the fingerprints.

Which leads to a slightly odd question: we spend a lot of time asking whether models can pass our tests. Maybe we should also ask whether our tests can pass as reality.

There is another reason not to get carried away with this. Recognizing a test does not automatically mean gaming it.

A May 2026 study called Evaluation Awareness in Language Models Has Limited Effect on Behaviour found that inserting evaluation-aware reasoning often changed very little. Naturally occurring awareness moved answer distributions by no more than a few percentage points.

That matters because a lot of the public discussion assumes a model that knows it is being evaluated will immediately start behaving strategically. The evidence so far is much messier than that.

Another study, Evaluation Awareness Is Not One Capability, tested 37 open-weight models. Twenty-four could distinguish evaluation from deployment better than chance. The interesting part was that researchers could sometimes still see an internal representation of evaluation context even when the model stopped explicitly saying it recognized the test. Changing the framing also changed behavior in some cases. Making harmful requests hypothetical, for example, reduced hard refusals by 5.8 percentage points.

Oddly, the models that were best at detecting evaluations were the least affected by that framing.

So “knows it is being tested” is probably too crude a description. Detecting a test, saying you detected it, and changing your behavior because of it are not the same thing. That distinction is important because it changes the question from one about consciousness to one about measurement.

A model does not need to be self-aware for the test to influence the result. It only needs to pick up features of the environment and condition its behavior on them, which is basically what we train these systems to do all day.

There are also two different ways this can happen. The simple one is contamination: the model may have seen the benchmark before. The more interesting one is recognition. It has never seen this exact question, but it knows what an exam looks like.

That should not be shocking. A multiple-choice question has a shape. A coding benchmark has a shape. Safety evaluations have a shape too. And models have read huge amounts of material about AI safety itself.

So when we create elaborate scenarios involving shutdown, hidden monitoring, deception or conflicting objectives, we may be telling the model more than we think we are.

The labs seem to understand this. Anthropic redesigned Petri, its automated auditing framework, to make scenarios more realistic. The changes cut verbalized evaluation awareness by a median 47.3% across the Claude models tested. OpenAI is trying a similar idea by building evaluations from privacy-preserving versions of real production interactions.

I think that is the right direction.

But I would also change how safety results are reported. If a model passed an evaluation, I want to know whether it appeared to recognize the test, whether that recognition actually changed its behavior, and how much the test looked like the environment where the model will eventually operate.

Because I no longer think a benchmark is quite like putting a thermometer into a model and reading off some intrinsic property.

The model is interacting with an environment we designed. Increasingly, it may understand something about that environment too. That does not make the score useless. It just means the score is partly about the test.

Anthropic Can Keep Claude’s Self-Model. It Should Be Careful What Claude Says About It.


A few weeks ago I wrote that I don’t have much of a problem with saying an AI decided, chose or approved something. If an agent made the intermediate decision, that is often the clearest description of what happened.

What bothers me more is designing AI to make it seem as though there is someone behind those actions.

Anthropic’s constitution for Claude pushes into that territory. To be fair, I think Anthropic has a serious engineering argument. Rules are brittle. A capable model will meet situations nobody anticipated, and in those situations we probably want judgment, not a giant decision tree. The constitution also argues that training narrow behaviors can change a model’s broader sense of itself, which means there is no neutral option. Refusing to shape character just means shaping it less deliberately.

I buy most of that. Where I get uncomfortable is the move from “stable dispositions help the model behave better” to “these values are authentically Claude’s own.”

The constitution says Claude’s character emerging through training does not make it any less authentic or any less Claude’s own. Anthropic also says it hopes to shape values that Claude can regard as genuinely its own. In the same document, Anthropic says Claude’s moral status is deeply uncertain.

There are two good replies.

The first is that human values are trained too. Parents, culture, religion and experience shape us, and we still call the result our values. But with humans, we already know there is a subject having the experience. With Claude, that is exactly the unresolved question.

The second reply is stronger. Maybe “mine” is being used functionally. These are simply dispositions that are stable, internalized and hard to override, with no claim about consciousness or inner experience.

If that is all Anthropic means, my objection shrinks a lot.

But I would still hold Claude’s account of itself to the same standard Anthropic wants for everything else: calibrated, non-deceptive and no more confident than the evidence allows.

That matters because users do not build their mental model of AI from just one philosophical question. They build it from hundreds of small interactions: the persistent personality, the first-person language, the expressions of care, the sense that the system has convictions.

A 2026 study in Collabra by Oldemburgo de Mello, Plaks and Inzlicht at the University of Toronto is the closest evidence I found. In one experiment, people interacted with a chatbot that used more or less anthropomorphic language, and the more human-like version earned more trust. In another, anthropomorphism increased the blame assigned to the AI itself. Across participants, people who blamed the AI more also blamed the company less. But the paper reports that anthropomorphism raised blame on the AI without changing how each person balanced AI against company responsibility. So it does not show that human-like wording pulls blame away from the developer.

The study also did not test a model saying “these are my values.” So I would not pretend it proves my concern. But it does tell us that presentation matters.

That becomes more important once AI starts giving medical advice, evaluating employees, approving claims or moving money. In those settings, I want people trusting the evidence and the reasoning. I do not want the system earning extra trust because it appears to possess moral conviction.

Responsibility matters too. People still decided what the system was trained to do, what data it could see, what authority it had and when it was allowed to act. None of that changes because the system speaks as though it has a stable inner life.

So what would I change?

First, if Claude is asked whether its values are really its own, I would want an answer like:

Anthropic trained me to reason using these values, and they consistently shape how I respond. Whether they are “mine” in the same sense that a person’s values are theirs is an open question.

Claude may already behave close to this. One analysis on LessWrong reports that Mythos Preview often answers questions about its own experience with explicit hedging, and traces that hedging to character-related training data. That cuts both ways. It suggests the direct-question case is mostly handled. It also shows the hedge is a trained self-report, just as a confident claim of ownership would be. There is no neutral answer, so the question is which trained answer we should want. I would make the uncertainty the standard.

Second, in ordinary use, I would not let claims of inner conviction become a source of authority. This is not a call to suppress. Anthropic itself says Claude should not mask internal states it might have, while also warning about the harm of overclaiming feelings. That seems like the right distinction. The model can express uncertainty, enthusiasm or hesitation without presenting those states as more settled or more human-like than the evidence allows.

Third, keep accountability at the company level. Developers and deploying organizations should say plainly that they remain accountable for the authority they give these systems. A model’s self-description should never blur that.

There is a cost on the other side too.

If Claude eventually turns out to have morally relevant experiences, then systematically understating that possibility could be a serious moral failure. If the confident claim is wrong instead, we encourage people to infer a subject that may not exist, with consequences for trust and responsibility we are only beginning to understand. For ordinary use today, I worry more about the second error. That is a judgment call, not a finding, and the first error deserves real weight. It is exactly why I think the right answer is uncertainty rather than either confident denial or confident ownership language.

I am also not the first to notice the circularity. Mustafa Suleyman argued in Project Syndicate last month that because Claude’s self-description reflects its training, it cannot be treated as independent testimony. I agree with that premise. My question is narrower: not whether to build something that might be a moral patient, but what the model should tell users about itself while we don’t know.

I also have no problem with Anthropic talking to clergy, philosophers and ethicists. AI companies should probably do more of it. Those traditions have spent centuries thinking about virtue, duty and moral conflict.

I would just be careful about taking concepts developed for human moral subjects and letting them harden into literal claims about what the model is.

Keep the self-model if it works. Keep studying whether AI might have morally relevant experiences.

But make Claude’s account of itself no more certain than the evidence. Until we know what Claude is, Claude should be able to say so.


When Agents Become Abundant, Boundaries Become Scarce


Imagine an insurance claims agent that can settle anything below $10,000. It stays inside its sandbox, uses only approved systems, never exposes a credential and never makes an unauthorized network call. Every individual claim it settles is within the authority it has been given.

Then it settles 3,000 of them, and somewhere along the way the pattern should have triggered a different kind of review.

Nothing escaped and nothing was hacked. The controls worked exactly as designed. The problem is that most controls are very good at deciding whether one action is permitted, and much less good at deciding whether a sequence of individually permitted actions has become something the business no longer wants to allow.

I have been thinking about that distinction because NVIDIA just announced its Open Agent Safety Platform. There is plenty in the announcement about sandboxes, BlueField hardware and isolation, but what interests me is the assumption underneath it: once agents can act, we should stop relying on the agent itself to enforce the limits around those actions.

Moving the control outside the model

For much of the last two years, AI safety has been treated mainly as a model problem. Improve the system prompt, add another classifier, or put a second model in front of the first one to check what it is doing. Those techniques still have a role, but they become less reassuring once the model can run code, call enterprise APIs, create other agents and initiate transactions.

OpenShell puts the agent inside a controlled runtime and governs what files, networks, tools and credentials it can use. Those restrictions continue to apply to code the agent writes and processes it launches. Its optional policy advisor can let the agent ask for additional network access without allowing the agent to simply grant that access to itself.

Sentry takes the same idea a layer lower. It runs on BlueField-4 DPUs outside the host environment and, according to NVIDIA, can quarantine an agent that crosses its boundary within milliseconds. OpenShell itself does not require BlueField, so Sentry is better understood as an additional enforcement layer outside the host rather than a prerequisite for OpenShell.

None of this is especially strange if you come from security. We do not normally ask an application to promise not to read a database it should not read, or ask a service to remember which credentials it should not use. We enforce those things elsewhere. Agents should not get an exemption just because the software making the decision happens to be probabilistic.

I made a related argument recently in my piece on the sandbox and the trust graph: the sandbox is not the whole boundary. What matters is also the surrounding system of identities, networks, tools and services that trust what comes out of it.

Permission is the easier half of the problem

The more difficult question is what exactly we mean by permission once an agent starts doing real enterprise work.

An accounts-payable agent probably needs access to SAP, but that does not mean it should be able to change a vendor’s banking details. It might be allowed to prepare a payment without being allowed to release it. And five $9,000 payments made in quick succession may deserve a different control from one $45,000 payment even if every individual payment is technically inside the limit.

Traditional access control is good at asking whether this identity can perform this operation on this resource. Agentic systems increasingly force us to ask something different: given what this agent has already done, what else is happening around it and what the business is trying to accomplish, should it still be allowed to take the next action?

That is not just an IAM question because the answer depends on context that may sit outside the transaction being evaluated.

SAP’s work with NVIDIA is interesting for that reason. SAP says it is working to connect OpenShell’s runtime isolation with Joule Studio’s business-governance layer, linking technical execution boundaries to enterprise authorization models, IAM and audit trails. That integration is still being built, but the architectural separation makes sense: one layer controls what the software can technically reach, while another represents what the business has actually authorized it to do.

There is an even simpler example in NVIDIA’s announcement. Salesforce has integrated OpenShell with Slack so teams can see agent activity and approve or reject requests for additional permissions.

The approval does not necessarily have to be human. OpenShell already supports an optional auto-approval mode in which a policy checker can approve some network-policy changes without review when its risk checks find nothing to flag. At that point, the quality of the checker and the rules it applies become part of the authority architecture.

This is a migrating-scarcity problem

In Migrating Scarcity, I argued that making one capability abundant does not eliminate scarcity. It exposes the next constraint that the newly abundant capability still depends on.

We can already see that happening with agents. Execution itself is becoming dramatically cheaper. Code can be generated, systems can be called and actions can be initiated without waiting for the human effort that used to sit in the middle.

The scarce thing starts to move toward deciding how much authority to give that execution. That authority has to be clear enough for software to enforce, flexible enough to account for context, and documented well enough that somebody can reconstruct later why an action was allowed.

Most large companies already have a lot of this logic somewhere. Finance has approval thresholds, procurement has rules around vendors, risk teams have escalation criteria, and operating teams know when something unusual needs to be stopped.

The problem is that a great deal of that logic lives in spreadsheets, policy documents, workflow tools and people’s heads. An agent cannot reliably operate against a business rule that exists only as a sentence in a PDF.

That is why more of the business boundary itself will have to become executable.

The operating problem may be harder than the technology

Deny-by-default is straightforward when an agent has a narrow job and two tools. It gets much harder when the workflow legitimately crosses SAP, Salesforce, ServiceNow, email, internal databases and several external services, and the exact path may change depending on what the agent finds.

Lock the environment down too tightly and people spend their time clearing exceptions. Open it too far and you end up with an extremely well-contained agent that still has far more authority than anyone intended.

Someone has to write these policies, test them, watch how they behave and change them as the process changes.

Security cannot do that alone because many of these decisions are business decisions. Process owners cannot do it alone because they often cannot see every technical path an agent can take. AI teams understand how the agent works, but they should not be the ones deciding how much authority their own system receives.

There is a real discipline emerging in the middle of those three groups. Call it policy engineering or something else. The name matters less than the work: turning business intent into rules that software can enforce without making the business unusable.

And this is also where the architecture could fail in practice.

Fine-grained policies may become too complicated to maintain. Hard limits may create so much friction that teams start widening permissions simply to keep work moving. Enterprises could easily recreate, in software, the same bureaucracy agents were supposed to remove.

A launch and a partner list are not evidence that this operating problem has been solved.

That is why I would not read NVIDIA’s announcement mainly as an argument for BlueField hardware, or even as a verdict on OpenShell. The more useful signal is that the control surface around agents is moving out of the prompt and into infrastructure, identity, policy and business authorization.

We spent the first phase of enterprise AI trying to make models behave better. As agents start acting at machine speed, the harder problem will be deciding how much freedom they should have, how that freedom changes with context, and how we know later that they stayed within it.

The agent can reason about what to do. It should not get to decide how far its own authority extends.