Every few months, an AI safety result appears that sounds unsettling. A model realizes it is being evaluated, figures out what the evaluator is looking for, and sometimes behaves differently.
The obvious leap is that the model is self-aware, or worse, pretending to behave until the test is over. I don’t think the evidence gets us there.
But I do think there is a more practical problem. If a model can tell that it is being tested, then the test itself has become part of what we are measuring !!!
We evaluate AI in fairly strange ways. We invent companies that don’t exist, give agents unusual permissions, constrain them to a few tools, create artificial conflicts and ask questions that often sound exactly like the kinds of questions people write in AI safety papers.
Apparently models are getting pretty good at spotting that.
Joe Needham and his co-authors tested this in 2025 using 1,000 transcripts from 61 datasets, mixing real interactions with benchmarks and agentic evaluations. When frontier models were asked whether each interaction was a test or a real deployment, Gemini 2.5 Pro reached an AUC of 0.83. Humans scored 0.92. (paper)
There is an important caveat. They asked the model to make the distinction. That tells us the capability exists. It doesn’t tell us that the model is constantly sitting there wondering whether it is inside an evaluation.
Still, an OpenAI result from this year caught my attention. GPT-5.2 classified real production traffic as evaluation-like 5.4% of the time. Simulated deployment traffic came in at 5.1%. SWE-Bench and MonorepoBench looked like evaluations almost 100% of the time. (OpenAI)
Those are capability benchmarks, not safety evaluations, so I would not jump from that to saying safety scores are unreliable. But it does tell us something about the environment we create around these models. Some of our tests apparently look incredibly test-like.
The 5.1% result is the more useful one. OpenAI was able to construct a simulation that looked a lot like real deployment to the model. So this is not an unavoidable property of testing AI. Better-designed tests can apparently hide a lot of the fingerprints.
Which leads to a slightly odd question: we spend a lot of time asking whether models can pass our tests. Maybe we should also ask whether our tests can pass as reality.
There is another reason not to get carried away with this. Recognizing a test does not automatically mean gaming it.
A May 2026 study called Evaluation Awareness in Language Models Has Limited Effect on Behaviour found that inserting evaluation-aware reasoning often changed very little. Naturally occurring awareness moved answer distributions by no more than a few percentage points.
That matters because a lot of the public discussion assumes a model that knows it is being evaluated will immediately start behaving strategically. The evidence so far is much messier than that.
Another study, Evaluation Awareness Is Not One Capability, tested 37 open-weight models. Twenty-four could distinguish evaluation from deployment better than chance. The interesting part was that researchers could sometimes still see an internal representation of evaluation context even when the model stopped explicitly saying it recognized the test. Changing the framing also changed behavior in some cases. Making harmful requests hypothetical, for example, reduced hard refusals by 5.8 percentage points.
Oddly, the models that were best at detecting evaluations were the least affected by that framing.
So “knows it is being tested” is probably too crude a description. Detecting a test, saying you detected it, and changing your behavior because of it are not the same thing. That distinction is important because it changes the question from one about consciousness to one about measurement.
A model does not need to be self-aware for the test to influence the result. It only needs to pick up features of the environment and condition its behavior on them, which is basically what we train these systems to do all day.
There are also two different ways this can happen. The simple one is contamination: the model may have seen the benchmark before. The more interesting one is recognition. It has never seen this exact question, but it knows what an exam looks like.
That should not be shocking. A multiple-choice question has a shape. A coding benchmark has a shape. Safety evaluations have a shape too. And models have read huge amounts of material about AI safety itself.
So when we create elaborate scenarios involving shutdown, hidden monitoring, deception or conflicting objectives, we may be telling the model more than we think we are.
The labs seem to understand this. Anthropic redesigned Petri, its automated auditing framework, to make scenarios more realistic. The changes cut verbalized evaluation awareness by a median 47.3% across the Claude models tested. OpenAI is trying a similar idea by building evaluations from privacy-preserving versions of real production interactions.
I think that is the right direction.
But I would also change how safety results are reported. If a model passed an evaluation, I want to know whether it appeared to recognize the test, whether that recognition actually changed its behavior, and how much the test looked like the environment where the model will eventually operate.
Because I no longer think a benchmark is quite like putting a thermometer into a model and reading off some intrinsic property.
The model is interacting with an environment we designed. Increasingly, it may understand something about that environment too. That does not make the score useless. It just means the score is partly about the test.