Intelligence Is Cheap. Permission Still Has to Be Built by Hand.


Toyota did not win the manufacturing wars of the 1980s because it had a better factory. American plants ran comparable equipment, comparable tolerances, often the same suppliers. Toyota won because it redesigned who had permission to act.

In the 1970s, Toyota gave line workers something most manufacturers would have considered reckless: the authority to stop the entire production line. The worker who pulled the andon cord wasn’t the most senior person in the building. They weren’t in a meeting with the plant manager. They were usually just the person standing closest to the defect, the one with the least formal power and the most immediate information. The andon cord was never a productivity tool. It was an authority architecture, a decision about who gets to act on what they see, without waiting for someone above them to see it too.

American manufacturers spent the better part of a decade copying the visible parts, the kanban cards, the quality circles, and mostly failed, because what they were copying wasn’t the actual advantage. The advantage was the redesigned permission underneath it, and permission doesn’t show up on a factory tour.


A Second Proof, Twenty Years Later

I’ve written before about Desktop Underwriter, the automated mortgage underwriting system Fannie Mae shipped in 1995. The value didn’t come from software that could evaluate a loan file faster than a human. Every competitor could eventually buy comparable software. The value came from what Fannie Mae attached to the system’s output: a waiver, relief from having to re-verify certain judgments the system had already made, provided a lender’s own data and documentation held up. That’s not automation. That’s an institution redesigning who could decide, who carried the risk when the decision was wrong, and how exceptions got handled, around a machine’s output.

Same mechanism, different industry, twenty years apart. The technology was necessary in both cases. It was never what got captured. What got captured was the permission architecture built around it, and I’ve called that capture Default Capture before: the winner is never whoever owns the technology, it’s whoever owns the constraint the technology creates downstream of itself.

AI is about to run the same test a third time.


The Question That Actually Matters

Most of the AI conversation happening in boardrooms right now is stuck on one question: can the model reason well enough to be trusted?

That’s the easy question, and the mortgage industry answered its version of it in 1995. The harder question, the one that actually determines whether an enterprise can act on what its systems produce, is different:

Can the organization allow the model to act?

Every technology transition creates a new abundance somewhere. The winners are never the organizations with the most of the newly abundant thing. They’re the ones that redesign permission around whatever became the constraint instead, before the market redesigns it for them, on someone else’s timeline.

That’s not a technology question. It’s the same question Toyota answered on a factory floor and Fannie Mae answered in a rulebook. Most enterprises haven’t answered it at all. They’ve bought the equivalent of the andon cord and left it bolted to the wall, unconnected to anything.


The Pilot Trap

Walk into almost any large enterprise right now and you’ll find the same three things: an enterprise AI license, a dozen point-solution pilots, and a Center of Excellence generating slide decks about theoretical time savings.

The demos are genuinely good. That was never the problem. The problem is that the company is putting a new engine into an old transmission. Every output still moves through the same approval chain built when the model needed supervision to be trustworthy. Every exception still follows the same escalation path designed for a system that used to be wrong a lot more often than it is now.

The model got dramatically better. The org chart didn’t get the memo. That gap is where the money goes to die, not in model cost, in the friction of an organization still authorizing decisions the way it did when authorization was the only safety mechanism available.

I don’t think this is a technology adoption problem. It’s an Organizational Rewiring Latency problem, and it’s a particularly nasty one, because unlike most of the shifts I’ve written about, this one doesn’t announce itself as a crisis. Nothing breaks. The pilots keep running. The demos keep landing well in the quarterly review. The company just quietly stays exactly as slow as it was before, with a much more expensive engine attached to the front of it.


The Blast Radius Problem

I’ve written before about the Law of Migrating Scarcity: when technology makes something abundant, value doesn’t disappear, it moves to whatever the abundance can’t dissolve. For decades, enterprise intelligence was that scarce resource, so companies built permission systems that assumed it would stay that way, slow, expensive, routed through whoever in the hierarchy had the most of it. Permission itself was never the scarce thing. It didn’t have to be. It only had to keep pace with a world where a decision moved as fast as the human hierarchy that had to bless it.

That world is ending. Models are commoditizing on schedule, the same way the underlying intelligence is, and the permission system built around its old scarcity is what’s left standing as the constraint. Permission isn’t becoming scarce out of nowhere. It’s being exposed as the bottleneck it was always going to become, the moment the thing it was rationing stopped being rare.

The advantage now belongs to whoever builds the clearest architecture for what a system is allowed to do without a person in the loop, and what still requires one. Within the decision and operational layers of that architecture, most companies still have exactly one lever: human review, on or off, applied uniformly regardless of stakes or track record. A usable version isn’t a single switch. It’s three tiers.

Tier 1, human-in-the-loop. The system recommends, a person decides before anything executes. High-consequence, low-frequency, hard-to-reverse decisions belong here, regulatory filings, large pricing exceptions, anything with real legal exposure attached.

Tier 2, human-on-the-loop. The system decides and acts, a person monitors and can intervene inside a defined window before the consequences compound. Most operational workflows land here once you’ve actually built some track record with the system.

Tier 3, human-out-of-the-loop, bounded. The system acts autonomously inside an explicit blast radius, a capped dollar amount, a reversible action, a pre-cleared category. A person audits the pattern, not the transaction.

Toyota never gave a worker unlimited authority. They could stop the line. They couldn’t redesign the factory, renegotiate with a supplier, or change the product roadmap. The authority had a blast radius, which is exactly what these tiers are designing. They’re a tool for two of the five layers I’ve written about before, decision permission and operational permission, not a replacement for the other three. A perfectly bounded Tier 3 can still fail if whoever audits it shares the same blind spot the system does, or if the outside world, regulators, customers, counterparties, never extends the market trust the internal architecture assumed it would have.

The actual design work isn’t picking a tier once and moving on. It’s the mechanism that promotes a decision from Tier 1 toward Tier 3 as the system earns trust in that specific domain, and demotes it the moment it doesn’t, the same kind of circuit breaker I’ve argued multi-agent systems need for cost, applied here to authority instead of spend. Almost nobody has built that mechanism.

Toyota’s cord and Fannie Mae’s rulebook both had an advantage this version doesn’t. A violation was visible. A worker could see a defect. An underwriting file either met the rule or it didn’t. A model that’s slowly drifting in quality doesn’t trip a wire, it just gets quietly worse, and the same review process built to catch it can end up sharing its blind spot, the exact failure I’ve written about in automated underwriting’s own verification layer. The promotion and demotion mechanism above only works if an organization can actually detect drift in a probabilistic system, and that detection problem is harder here than it ever was on a factory floor. Building a Tier 3 that looks bounded on paper without solving that problem first is how the blast radius stops meaning anything.

The reason this becomes a durable advantage, and not just a nice-to-have, is that permission architectures are hard to copy. A competitor can buy the same model. They can license the same software. They can hire the same consultants who hired the same consultants. What they can’t buy off the shelf is the accumulated trust, the operating data, and the specific decision boundaries that let one organization move with confidence while another is still in a meeting arguing about who has the authority to approve the meeting’s outcome.


What Autonomy Actually Costs When You Get It Wrong

It would be dishonest to make this argument without naming what it costs when it’s done badly. Expanding autonomy without a real governance mechanism doesn’t remove risk, it just changes its shape, from slow and visible to fast and compounding. A bad approval chain produces one bad decision at a time, and someone downstream usually catches it. A badly bounded Tier 3 produces the same bad decision at machine speed, thousands of times, before anyone notices anything’s wrong.

That’s not an argument against building this. It’s the argument for building it on purpose instead of drifting into it. The companies that get hurt in this transition won’t be the ones that moved too slowly on autonomy. They’ll be the ones that expanded it without a defined blast radius, without an audit trail, and without a name attached to who owns it when it fails. Tiering, bounding, and auditing is the whole difference between deliberate autonomy and an accident waiting for a headline.


Three Questions Worth Asking This Quarter

If you’re the one accountable for this inside your company, the diagnostic isn’t how many pilots you’re running. It’s:

Where does a decision still require a human sign-off purely out of habit, not because the stakes or the risk actually call for it?

For your highest-volume automated processes, do you have a defined blast radius and an audit mechanism, or just an on/off switch?

Who owns the outcome when a system acts on its own and gets it wrong, and do they know yet that it’s their job?

Toyota moved authority from headquarters to the factory floor. Mortgage underwriting moved authority from individual judgment to an institutional system. AI is going to move authority from human execution to bounded autonomous systems. The technology changes each time. The pattern underneath it doesn’t.

Most companies can’t answer the second or third question today. That gap, not model capability, will separate the companies that become the default this decade from the ones that spend it running increasingly sophisticated pilots.

Where scarcity goes when it leaves


A few weeks ago I wrote that boards are treating AI intelligence as a permanently scarce, permanently expensive input, and that this assumption is already cracking. A few people asked me a fair follow up question. If intelligence stops being the scarce thing, where does the scarcity actually go. It doesn’t just vanish. Someone always ends up holding it.

I have watched this happen up close twice in my own career, and I know a third example only from history, but it is the cleanest one to start with because you can actually see the migration happen.

The first time was containers. Malcolm McLean did not invent a faster ship. He invented a standard box, and that box made loading and unloading cargo dramatically cheaper. Everyone assumed the story was over once ships stopped idling in port for a week at a time. It wasn’t. Ports suddenly needed acres of land to stack containers. Rail lines had to sync up with ship schedules. Crane operators and terminal planners became more valuable than the stevedores whose jobs the container had just eliminated. The scarcity did not disappear when loading got cheap. It walked a few hundred yards down the dock and set up shop in land, rail and coordination.

The second time was the shift from FTE pricing to outcome pricing in IT and BPO, which is a conversation I have had more times than I can count over the last two years. For a long time, the constraint looked like headcount. You needed bodies to run processes, so you priced by the body. As automation and now agentic AI made raw execution cheaper, everyone assumed the constraint would just dissolve along with the headcount. It didn’t, because the bottleneck was never just labor. It was the entire system built around buying, measuring and governing labor. Procurement teams that are set up to negotiate FTE contracts are, frankly, not set up to negotiate outcome contracts, and it isn’t only because procurement is slow. When a workflow runs through a mix of human teams, vendor agents and enterprise software, agreeing on who actually caused a given outcome is a genuinely messy problem, not just a paperwork one. Finance teams that know how to forecast a headcount ramp don’t automatically know how to forecast a variable outcome fee they can’t cleanly attribute in the first place. That system, not the labor it was built to manage, is the thing that is actually scarce right now. I said back in February that this is exactly why Khosla’s five year timeline for IT and BPO extinction won’t hold. Enterprises are slow to redesign the muscle that buys and governs work, and rebuilding that muscle takes a lot longer than swapping the underlying technology.

The third time is AI, and we are living through the early innings of it.

The public conversation is entirely about model capability. Whose benchmark is better this month, whose inference is cheaper, whose context window is longer. That is the visible layer, and it is genuinely moving fast. But if you sit in enough steering committee meetings, as I do, you notice the real conversation has already shifted somewhere else. Nobody is asking whether the model is good enough anymore. They are asking who is accountable when an agent acts on its own, how you audit a decision a model made six tool calls deep, and whether legal and risk can sign off before the business quarter ends. Legal, risk and compliance are quietly becoming the functions that decide how fast AI actually ships, not engineering. And this isn’t only a soft, organizational story either. I wrote back in February that Jevons paradox is still very much alive in AI, cheap intelligence doesn’t shrink total demand, it multiplies the number of things people try to do with it. That multiplication is what’s straining power grids and chip supply right now, and it will keep straining them. The bottleneck isn’t only moving into legal’s inbox. It’s splitting, part of it lands on governance, part of it lands on the physical world’s ability to keep up.

That is the pattern, and it holds across all three examples. When something that used to be the bottleneck becomes cheap, the constraint does not evaporate. It relocates to whatever has to absorb the new abundance. Land and rail after containers. Contracts and procurement after outcome based pricing. Governance and organizational readiness after intelligence.

I want to be honest about where this framework is weaker than it sounds. It is easy to find three examples that fit a pattern after the fact. The real test is whether it predicts anything, and whether there are cases where it breaks. It does break sometimes. The cloud is actually the interesting counterexample here, not the confirming one. When compute got cheap, the constraint should have moved to independent architecture and security specialists. Instead, the hyperscalers largely built and sold that layer themselves. AWS didn’t watch a market of third party cloud governance firms spring up and capture the value, it built Control Tower and Security Hub and kept the margin in house. The incumbents who already controlled the abundant layer often reach up and grab the scarce layer too, they just do it slower than a scrappy new entrant would. Call it the adjacent ownership problem if you want a name for it.

But I don’t think AI plays out quite the same way, and it’s worth being precise about why. AWS could absorb cloud security because cloud security is still fundamentally tooling, dashboards, policies, audit logs, things a vendor can build and sell. Frontier labs can absorb model guardrails the same way, and several of them are trying to. What they cannot absorb is the thing sitting underneath the tooling: whose name is on the regulatory filing, who eats the liability when an agent makes a bad call, who signs the indemnity clause. That layer doesn’t move to the vendor no matter how good their safety tooling gets. It stays inside the enterprise. So the AI version of this migration may actually be more durable than the cloud version, the technical guardrails can be commoditized by whoever owns the model, but the accountability cannot be outsourced the same way. Though I’d bet even that has a shelf life. The moment someone figures out how to price and package agentic risk the way insurers price everything else, that liability becomes securitizable too, and the scarce resource quietly becomes actuarial expertise instead. Scarcity doesn’t stop migrating just because it hit an enterprise’s balance sheet. And I should be honest that this bottleneck doesn’t always slow things down the tidy way a land shortage slows down a port. Sometimes a business unit just routes around legal entirely, the way shadow IT always found a way around IT, and what looks like delay from the boardroom is actually unowned risk quietly accumulating somewhere nobody is tracking it yet.

So if you are trying to figure out where value is actually migrating in your own AI strategy, the technology roadmap is the least useful thing to stare at. Watch where the friction is showing up instead. Watch which meetings in your company have gotten longer, not shorter, since AI arrived. Watch which job titles didn’t exist eighteen months ago and are now impossible to hire for fast enough. Watch whether your procurement team can even write a contract for an outcome nobody has priced before.

The technology tells you what just became possible. The friction tells you where value is about to accumulate. And the companies that win the next few years will not be the ones who called the breakthrough early. They will be the ones who noticed where the scarcity went after it left.

The First Permission Architecture


How Automated Underwriting Revealed the Real Bottleneck in Autonomous Systems

The hardest problem with autonomous systems has never been only intelligence. It’s permission. A system can produce a decision in milliseconds. The harder question, the one that actually determines whether an organization can act on that decision, is whether anyone has redesigned liability, workflow, and trust around it.

The mortgage industry answered that question in 1995. Then it stopped asking it carefully enough, and broke the answer thirteen years later in a way worth studying as closely as the original solution.

The system

In October 1994, Fannie Mae piloted a program called Desktop Underwriter. By June 1995 it was live in production. DU was what the AI field of that era called an expert system: not a model that learned patterns from data the way modern machine learning does, but a rules engine that encoded the judgment of experienced human underwriters into a structured decision process, built to evaluate incomplete, unverified, and sometimes conflicting borrower data against a structured set of underwriting rules.

Feed it a loan file, and DU would return a credit recommendation and an eligibility recommendation, an automated judgment on whether a mortgage met Fannie Mae’s standards, in minutes instead of the days a manual file review took. Freddie Mac shipped a competing system, Loan Prospector, around the same time.

The consequential design choice wasn’t the automation. It was what came attached to the recommendation. When a loan file received DU’s top designation, Approve/Eligible, Fannie Mae extended lenders a waiver: relief from having to represent and warrant that the loan met Fannie Mae’s underwriting and eligibility standards, provided the lender’s data was accurate and properly documented. If DU said yes, and the paperwork behind that yes checked out, Fannie Mae absorbed a defined slice of the underwriting risk, the part tied to credit and eligibility judgment. The lender still carried the rest: data accuracy, fraud, documentation, the obligations no waiver ever touched.

That’s not automation. That’s a liability transfer, conditioned on a machine’s output, bounded by data-integrity rules the lender had to satisfy to earn it. Nobody at Fannie Mae in 1995 would have called it this, but it’s what I mean by an offensive permission architecture: not just letting a system make recommendations, but redesigning liability, workflows, and incentives around what it recommends, so the organization can act on the machine’s judgment without waiting for a human to re-verify every step of it. I use “offensive” deliberately and narrowly here: not aggressive deployment without safeguards, but an organization building the permission to act before the market or a regulator forces it to ask for that permission later, on someone else’s terms.

The advantage was never going to belong to whoever owned the automated underwriting system. Fannie Mae made DU available to every lender on identical terms. It was going to belong to whoever rebuilt their institution around the new source of authority DU created.

The permission stack

Pull the case apart and it separates cleanly into five questions, and the mortgage industry had to answer all five before automated underwriting became a source of advantage rather than a novelty.

Decision permission: can the system’s output count as a real decision, not just an input to one? DU cleared this by producing a recommendation lenders could act on directly.

Liability permission: when the decision is wrong, who absorbs it? Fannie Mae did, on Approve/Eligible loans that met the data-integrity conditions. Without this layer, DU is just a faster opinion. With it, DU is authority.

Operational permission: can the organization actually execute at the standard the authority requires? A lender had to have clean data pipelines and documentation discipline strong enough to survive scrutiny, or the waiver didn’t apply. This is the layer most companies underbuild, because it’s unglamorous compared to the technology itself.

Verification permission: does the organization’s own check on the system actually catch a systemic failure, or does it just confirm the system complied with its own rules? Fannie Mae required lenders to run post-closing quality control review on a sample of closed loans, checking that DU’s findings were properly resolved and documented. That’s a real verification layer, and for years it looked sufficient.

Verification is the layer most likely to create false confidence. A review process can be rigorous, sampled files, documented findings, signed-off reviews, and still be structurally blind, if the reviewer and the system share the same underlying assumptions. Fannie Mae’s QC reviewers were checking loans against the same underwriting guidelines DU had encoded, not independently assessing whether those guidelines still matched reality. It wasn’t that nobody was checking. Lenders sampled files constantly, documented every finding, signed off on schedule. The checking simply couldn’t see past a blind spot it shared with the thing it was checking. When Fannie Mae and Freddie Mac themselves loosened those guidelines through the 2000s, to accept lower credit scores, less documentation, less money down, the QC process could still confirm compliance. It just meant less, because the rulebook it was checking against had moved.

Market permission: will the parties outside the transaction, investors buying the resulting mortgage-backed securities, regulators overseeing the system, trust an outcome a machine helped produce? This was the layer verification was supposed to protect. It took the industry years to earn, and it was the layer that failed most visibly when trust collapsed in 2008.

These five layers don’t sit on top of each other so much as feed into each other. Faster decision permission creates operational pressure. Operational pressure, left unchecked, is what quietly erodes verification. Eroded verification is what eventually costs an institution its market permission. Pull on any one layer and the others move.

Most conversations about AI autonomy right now are stuck entirely on the first layer, whether the system can reason well enough to be trusted with a decision. The mortgage industry’s experience says that’s the easy one. The remaining four layers are where the actual advantage, and the actual risk, live.

The capture

Lenders who built the operational permission to qualify for the waiver, clean data pipelines, documentation that would hold up, systems that could feed DU reliably, got faster closings and lower risk retention than lenders who didn’t. That’s not a marginal efficiency gain. Between the first and second halves of the 1990s, the volume of mortgage-backed securities issued and guaranteed by Fannie Mae and Freddie Mac combined jumped from roughly $127 billion to $314 billion. Automated underwriting wasn’t the only driver of that, but it was a structural one: it changed what speed and scale were possible for lenders who’d built around it.

Notice what scarcity actually did here, because it’s the whole argument in one case. Before DU, one of the scarce resources was underwriting judgment itself, and it lived inside individual underwriters who couldn’t be copied or scaled. DU made that judgment portable and cheap. The scarce resource didn’t vanish, it moved, to whichever lender had built the institutional machinery to trust an automated decision enough to act on it at scale. That’s Default Capture: the winner isn’t whoever owns the technology, it’s whoever owns the constraint the technology creates downstream of itself, and builds the permission stack to act on that constraint before anyone else does.

The break

Here’s the part a triumphant case study would leave out.

The waiver was built for a specific kind of loan file: standard documentation, standard verification, standard borrower profiles, priced against risk assumptions Fannie Mae’s underwriters had spent decades calibrating. What changed through the 2000s wasn’t that riskier files quietly slipped through unnoticed. Fannie Mae and Freddie Mac actively loosened the rules DU was checking against, under real competitive pressure. Private-label securitizers were taking market share fast, the GSEs’ combined share of new mortgages fell from 52 percent in 2002 to 44 percent by 2006, and both agencies responded by expanding what counted as an acceptable loan: lower credit thresholds, zero-down products, wider acceptance of low and no-documentation files. Fannie Mae authorized more than eleven thousand underwriting variances in 2005 alone. The authorization architecture didn’t fail by drifting out of sync with a changing world. It failed because the people who owned the rulebook kept rewriting it to chase volume.

The lesson isn’t that automated underwriting caused the 2008 crisis. It didn’t, on its own, and plenty of other forces did more damage: private-label securitization, layered risk, ratings failures among them. The lesson is less comfortable than a single cause would be. The same architecture that safely accelerated standardized underwriting for a decade could be pointed at a much riskier target once the people who controlled it chose to loosen what the permission covered. The system didn’t fail by breaking. It failed by continuing to work exactly as designed, on a rulebook its own owners had rewritten to permit exactly the inputs it was never supposed to see.

When the crisis forced a reckoning, permission got rebuilt from the outside. In September 2012, the Federal Housing Finance Agency, then overseeing Fannie Mae and Freddie Mac in conservatorship, directed both GSEs to adopt a new representation and warranty framework, one with harder, more codified rules about when relief applied, tied to specific payment-history performance rather than a point-in-time automated recommendation alone. That framework wasn’t something any single lender could engineer its way into faster than competitors. It was imposed, uniformly, by a regulator, as a condition of an industry that had lost the market’s trust. Market permission, the layer nobody had been watching closely, was the one that failed most visibly, and once it did, the other four layers couldn’t keep operating at the same scale, even where they still functioned fine on their own terms.

Why the distinction matters

This is close to the cleanest real-world illustration of a distinction I keep coming back to: internal permission and external permission are different problems with different playbooks.

Fannie Mae’s original DU waiver was internal permission. It was a counterparty relationship, Fannie Mae and an individual lender, that a lender could earn through its own engineering: better data, better documentation, better systems. That’s a problem a company can solve on its own timeline, and the lenders who solved it early captured real advantage before competitors caught up.

The 2012 rep and warranty framework was external permission. It came from a federal regulator, after a crisis, uniformly, on a timeline no single company controlled. No amount of internal engineering discipline would have gotten a lender there faster. The authority to shape that outcome sat with FHFA, not with the industry.

Conflating those two, treating a regulator-controlled reset as something you can outbuild the way you outbuild a competitor’s data pipeline, is the exact category error I’d warn anyone in a regulated industry against making.

The lesson for what comes next

Strip away the paper and the decade, and the mechanism is identical to what enterprises are building around AI agents right now: bounded automated authority, a defined answer to who’s liable when the system is wrong, and a real, capturable advantage for whoever engineers the full permission stack before their competitors do.

Lay the two eras side by side and the mapping is almost exact. DU’s Approve/Eligible recommendation then is an agent approving a transaction or shipping code now. Fannie Mae’s rep-and-warranty waiver then is whoever indemnifies when the agent is wrong now. Clean data pipelines and documentation discipline then are sandboxing, API boundaries, and execution limits now. Post-closing QC review then is test suites and human-in-the-loop spot checks now. Trust from MBS investors and regulators then is regulatory clearance and user trust in the output now. Same five questions, thirty years apart.

Which companies actually get to run their agents at scale won’t be settled by model quality alone, the same way it wasn’t settled by whose rules engine evaluated files best in 1995. It’ll be settled by who builds the liability, operational, verification, and market permission underneath the decision, on their own timeline, before a regulator builds it for them on someone else’s. Intelligence is getting cheap fast. Permission still has to be built by hand, one institution at a time.