Hyperautomation Doesn’t Need Superintelligence


I’m increasingly skeptical of the assumption that progress in AI reasoning puts us on a path to superintelligence.

The expectation of far more powerful AI isn’t imaginary. In his October 2024 essay Machines of Loving Grace, Anthropic’s Dario Amodei described the possibility of a “country of geniuses in a datacenter.” He thought such powerful AI could arrive as early as 2026, though he acknowledged it might take considerably longer.

I can understand the optimism. Reasoning models have improved enormously. They solve difficult mathematical problems, write and debug code, and find solutions that would take human experts considerable time. Give them more opportunities to search, test and refine their answers, and performance often improves.

I expect that progress to continue. What I question is whether extending today’s methods gets us to machines that substantially exceed the best human minds across a broad range of intellectual tasks.

I think we’re heading toward something enormously valuable anyway: hyperautomation. I use the term in its established sense of combining AI with conventional automation to handle increasingly complex business processes. The economic case for doing this doesn’t depend on superintelligence.

Search needs a scorekeeper

A significant part of recent progress has come from reinforcement learning and additional computation devoted to finding better answers. Models generate possible solutions, receive feedback and improve.

DeepSeek-R1, subsequently published in Nature, is a good example. Its researchers used rule-based rewards for verifiable reasoning tasks, avoiding neural reward models in that setting because of their vulnerability to reward hacking during large-scale reinforcement learning.

This works remarkably well when evaluation is reliable. Mathematical answers can be checked. Code can be compiled and tested. Algorithms can be measured against known objectives.

But code that passes its tests may still solve the wrong problem. An algorithm can optimize its assigned metric while making the overall system worse.

Another model can judge the output, but that introduces its own difficulties. Research by Gao, Schulman and Hilton demonstrated that pushing optimization too aggressively against an imperfect reward model can eventually degrade performance against the underlying objective.

The harder problem really isn’t generating more candidate answers – rather it’s in figuring out which ones are good when the criteria are uncertain.

But AI is already making discoveries

The counterexamples deserve serious attention.

AlphaFold made an extraordinary advance in predicting protein structures. It learned from experimentally determined structures, giving it a foundation of physical evidence from which to develop predictions. Its confidence estimates are useful, although they don’t guarantee correctness.

AlphaEvolve is even more interesting. Google’s system found a method to multiply 4×4 complex matrices using 48 scalar multiplications, improving on a longstanding result. It also designed components of a gradient-based optimization procedure used to discover matrix multiplication algorithms.

These are genuine advances, and I absolutely expect many more. I am quite excited about what else will come out from this direction !

But AlphaEvolve also illustrates the importance of evaluation. It generates programs and tests them against measurable objectives. Google describes the system as applicable to problems whose solutions can be expressed as algorithms and automatically verified.

That’s a powerful method for discovery. It doesn’t establish that the system can independently recognize when an entire scientific framework is inadequate, develop an alternative and determine how to validate it.

Humans don’t begin with perfect scorekeepers either. Scientists invent experiments precisely because existing evidence is insufficient. AI may eventually become exceptionally good at this. I just haven’t seen convincing evidence that the current approach reliably generalizes to such problems.

Berkeley’s A-Lab offers another example. Its autonomous laboratory demonstrated impressive materials synthesis, but subsequent scrutiny raised questions about material identification and novelty. In a January 2026 correction, the authors clarified that the materials described as new were new to their prediction platform, not necessarily new to science. Their reanalysis confirmed 36 of 40 reported successes, with four inconclusive.

The automation was real. Establishing exactly what had been discovered required additional scrutiny.

Hyperautomation is already a big enough prize

This distinction matters in enterprise AI, where I spend much of my time.

Consider accounts payable. An agent can extract invoice data, match purchase orders, check receipts, apply payment policies, identify exceptions and initiate approved payments. Much of this works because the checks are deterministic. The invoice matches or it doesn’t. The payment falls within policy or it doesn’t. The transaction completes or fails.

The important design question is often not how intelligent the agent is, but what tells it when it’s wrong.

Move into a disputed invoice involving a strategic supplier, and the evaluation becomes less straightforward. Should the company enforce the contract, preserve the relationship or renegotiate the terms? An AI system may eventually make many of those decisions too, but it needs some way to resolve competing objectives and establish its authority to act.

We can automate an extraordinary amount of work without solving every such problem. Claims processing, financial close, customer service and software development can all benefit from combinations of reasoning models, conventional software, classifiers and human oversight. Not every task needs an LLM, much less a superintelligent one.

There are difficult engineering and organizational problems to solve. Integration, security, context, evaluation and process redesign remain substantial challenges. But addressing them could transform enterprise economics even if superintelligence never arrives.

This fits the argument I made in Migrating Scarcity. As execution becomes cheaper, the constraint moves toward specifying objectives, evaluating outcomes and deciding who or what has authority to act. Better automation doesn’t eliminate those questions. It makes them more important.

What would change my mind?

I’d take the current path to superintelligence more seriously if AI systems began independently identifying important scientific problems, constructing new explanations and designing experiments whose novel predictions were subsequently confirmed by independent researchers.

I’d also look for systems that recognize flaws in their own evaluation methods and develop better ones. A meaningful demonstration would be an AI discovering that its existing test rewarded the wrong behavior, designing a replacement, and showing that the replacement predicts independently measured outcomes more accurately on unfamiliar problems. Repeating that across different domains would be much stronger evidence than improvement on another benchmark.

I don’t know whether superintelligence is possible or whether some extension of today’s methods will eventually produce it. Scaling has surprised us before. But possibility isn’t evidence, and progress in search and automation doesn’t by itself establish a path to superintelligence.

I think hyperautomation will change the economics of entire industries. We have barely begun capturing that opportunity, and the engineering work ahead is substantial.

We don’t need to call it superintelligence to recognize how valuable it could become.

Your AI Safety Score May Be Measuring the Test, Too


Every few months, an AI safety result appears that sounds unsettling. A model realizes it is being evaluated, figures out what the evaluator is looking for, and sometimes behaves differently.

The obvious leap is that the model is self-aware, or worse, pretending to behave until the test is over. I don’t think the evidence gets us there.

But I do think there is a more practical problem. If a model can tell that it is being tested, then the test itself has become part of what we are measuring !!!

We evaluate AI in fairly strange ways. We invent companies that don’t exist, give agents unusual permissions, constrain them to a few tools, create artificial conflicts and ask questions that often sound exactly like the kinds of questions people write in AI safety papers.

Apparently models are getting pretty good at spotting that.

Joe Needham and his co-authors tested this in 2025 using 1,000 transcripts from 61 datasets, mixing real interactions with benchmarks and agentic evaluations. When frontier models were asked whether each interaction was a test or a real deployment, Gemini 2.5 Pro reached an AUC of 0.83. Humans scored 0.92. (paper)

There is an important caveat. They asked the model to make the distinction. That tells us the capability exists. It doesn’t tell us that the model is constantly sitting there wondering whether it is inside an evaluation.

Still, an OpenAI result from this year caught my attention. GPT-5.2 classified real production traffic as evaluation-like 5.4% of the time. Simulated deployment traffic came in at 5.1%. SWE-Bench and MonorepoBench looked like evaluations almost 100% of the time. (OpenAI)

Those are capability benchmarks, not safety evaluations, so I would not jump from that to saying safety scores are unreliable. But it does tell us something about the environment we create around these models. Some of our tests apparently look incredibly test-like.

The 5.1% result is the more useful one. OpenAI was able to construct a simulation that looked a lot like real deployment to the model. So this is not an unavoidable property of testing AI. Better-designed tests can apparently hide a lot of the fingerprints.

Which leads to a slightly odd question: we spend a lot of time asking whether models can pass our tests. Maybe we should also ask whether our tests can pass as reality.

There is another reason not to get carried away with this. Recognizing a test does not automatically mean gaming it.

A May 2026 study called Evaluation Awareness in Language Models Has Limited Effect on Behaviour found that inserting evaluation-aware reasoning often changed very little. Naturally occurring awareness moved answer distributions by no more than a few percentage points.

That matters because a lot of the public discussion assumes a model that knows it is being evaluated will immediately start behaving strategically. The evidence so far is much messier than that.

Another study, Evaluation Awareness Is Not One Capability, tested 37 open-weight models. Twenty-four could distinguish evaluation from deployment better than chance. The interesting part was that researchers could sometimes still see an internal representation of evaluation context even when the model stopped explicitly saying it recognized the test. Changing the framing also changed behavior in some cases. Making harmful requests hypothetical, for example, reduced hard refusals by 5.8 percentage points.

Oddly, the models that were best at detecting evaluations were the least affected by that framing.

So “knows it is being tested” is probably too crude a description. Detecting a test, saying you detected it, and changing your behavior because of it are not the same thing. That distinction is important because it changes the question from one about consciousness to one about measurement.

A model does not need to be self-aware for the test to influence the result. It only needs to pick up features of the environment and condition its behavior on them, which is basically what we train these systems to do all day.

There are also two different ways this can happen. The simple one is contamination: the model may have seen the benchmark before. The more interesting one is recognition. It has never seen this exact question, but it knows what an exam looks like.

That should not be shocking. A multiple-choice question has a shape. A coding benchmark has a shape. Safety evaluations have a shape too. And models have read huge amounts of material about AI safety itself.

So when we create elaborate scenarios involving shutdown, hidden monitoring, deception or conflicting objectives, we may be telling the model more than we think we are.

The labs seem to understand this. Anthropic redesigned Petri, its automated auditing framework, to make scenarios more realistic. The changes cut verbalized evaluation awareness by a median 47.3% across the Claude models tested. OpenAI is trying a similar idea by building evaluations from privacy-preserving versions of real production interactions.

I think that is the right direction.

But I would also change how safety results are reported. If a model passed an evaluation, I want to know whether it appeared to recognize the test, whether that recognition actually changed its behavior, and how much the test looked like the environment where the model will eventually operate.

Because I no longer think a benchmark is quite like putting a thermometer into a model and reading off some intrinsic property.

The model is interacting with an environment we designed. Increasingly, it may understand something about that environment too. That does not make the score useless. It just means the score is partly about the test.

Anthropic Can Keep Claude’s Self-Model. It Should Be Careful What Claude Says About It.


A few weeks ago I wrote that I don’t have much of a problem with saying an AI decided, chose or approved something. If an agent made the intermediate decision, that is often the clearest description of what happened.

What bothers me more is designing AI to make it seem as though there is someone behind those actions.

Anthropic’s constitution for Claude pushes into that territory. To be fair, I think Anthropic has a serious engineering argument. Rules are brittle. A capable model will meet situations nobody anticipated, and in those situations we probably want judgment, not a giant decision tree. The constitution also argues that training narrow behaviors can change a model’s broader sense of itself, which means there is no neutral option. Refusing to shape character just means shaping it less deliberately.

I buy most of that. Where I get uncomfortable is the move from “stable dispositions help the model behave better” to “these values are authentically Claude’s own.”

The constitution says Claude’s character emerging through training does not make it any less authentic or any less Claude’s own. Anthropic also says it hopes to shape values that Claude can regard as genuinely its own. In the same document, Anthropic says Claude’s moral status is deeply uncertain.

There are two good replies.

The first is that human values are trained too. Parents, culture, religion and experience shape us, and we still call the result our values. But with humans, we already know there is a subject having the experience. With Claude, that is exactly the unresolved question.

The second reply is stronger. Maybe “mine” is being used functionally. These are simply dispositions that are stable, internalized and hard to override, with no claim about consciousness or inner experience.

If that is all Anthropic means, my objection shrinks a lot.

But I would still hold Claude’s account of itself to the same standard Anthropic wants for everything else: calibrated, non-deceptive and no more confident than the evidence allows.

That matters because users do not build their mental model of AI from just one philosophical question. They build it from hundreds of small interactions: the persistent personality, the first-person language, the expressions of care, the sense that the system has convictions.

A 2026 study in Collabra by Oldemburgo de Mello, Plaks and Inzlicht at the University of Toronto is the closest evidence I found. In one experiment, people interacted with a chatbot that used more or less anthropomorphic language, and the more human-like version earned more trust. In another, anthropomorphism increased the blame assigned to the AI itself. Across participants, people who blamed the AI more also blamed the company less. But the paper reports that anthropomorphism raised blame on the AI without changing how each person balanced AI against company responsibility. So it does not show that human-like wording pulls blame away from the developer.

The study also did not test a model saying “these are my values.” So I would not pretend it proves my concern. But it does tell us that presentation matters.

That becomes more important once AI starts giving medical advice, evaluating employees, approving claims or moving money. In those settings, I want people trusting the evidence and the reasoning. I do not want the system earning extra trust because it appears to possess moral conviction.

Responsibility matters too. People still decided what the system was trained to do, what data it could see, what authority it had and when it was allowed to act. None of that changes because the system speaks as though it has a stable inner life.

So what would I change?

First, if Claude is asked whether its values are really its own, I would want an answer like:

Anthropic trained me to reason using these values, and they consistently shape how I respond. Whether they are “mine” in the same sense that a person’s values are theirs is an open question.

Claude may already behave close to this. One analysis on LessWrong reports that Mythos Preview often answers questions about its own experience with explicit hedging, and traces that hedging to character-related training data. That cuts both ways. It suggests the direct-question case is mostly handled. It also shows the hedge is a trained self-report, just as a confident claim of ownership would be. There is no neutral answer, so the question is which trained answer we should want. I would make the uncertainty the standard.

Second, in ordinary use, I would not let claims of inner conviction become a source of authority. This is not a call to suppress. Anthropic itself says Claude should not mask internal states it might have, while also warning about the harm of overclaiming feelings. That seems like the right distinction. The model can express uncertainty, enthusiasm or hesitation without presenting those states as more settled or more human-like than the evidence allows.

Third, keep accountability at the company level. Developers and deploying organizations should say plainly that they remain accountable for the authority they give these systems. A model’s self-description should never blur that.

There is a cost on the other side too.

If Claude eventually turns out to have morally relevant experiences, then systematically understating that possibility could be a serious moral failure. If the confident claim is wrong instead, we encourage people to infer a subject that may not exist, with consequences for trust and responsibility we are only beginning to understand. For ordinary use today, I worry more about the second error. That is a judgment call, not a finding, and the first error deserves real weight. It is exactly why I think the right answer is uncertainty rather than either confident denial or confident ownership language.

I am also not the first to notice the circularity. Mustafa Suleyman argued in Project Syndicate last month that because Claude’s self-description reflects its training, it cannot be treated as independent testimony. I agree with that premise. My question is narrower: not whether to build something that might be a moral patient, but what the model should tell users about itself while we don’t know.

I also have no problem with Anthropic talking to clergy, philosophers and ethicists. AI companies should probably do more of it. Those traditions have spent centuries thinking about virtue, duty and moral conflict.

I would just be careful about taking concepts developed for human moral subjects and letting them harden into literal claims about what the model is.

Keep the self-model if it works. Keep studying whether AI might have morally relevant experiences.

But make Claude’s account of itself no more certain than the evidence. Until we know what Claude is, Claude should be able to say so.