A call center manager pulls up the dashboard the way she does most mornings, coffee still too hot to drink. Automation rate: up. Cost per contact: down. Service level: the best it’s been all quarter. Every number on the screen is green, and every number on the screen is telling her the year is going well.
So why does the floor feel harder to run than it did a year ago?
She can’t point to anything specific. Nobody’s quit in weeks. Escalations aren’t up, exactly. It’s just that the calls that make it to her agents lately seem to take more out of them – more silence on the line while someone reads a note, more calls that get transferred twice before they’re solved, more agents stopping by her desk to ask what they’re allowed to do about a situation the script doesn’t cover. None of that shows up as a number. The dashboard doesn’t have a field for it.
Here’s what she’s actually looking at, if she wrote it out.
Before AI, her team handled all 100 cases that came in on a given day. Some were easy – a password reset, a billing question with an obvious answer. Some were hard- an angry customer, three conflicting records, a policy that didn’t fit the situation. But on average, the work was representative of the whole. A typical case was, roughly, a real case.
Suppose AI now handles the easiest 70.
The 30 that reach her agents aren’t a smaller version of the old 100. They’re a different population. They contain more ambiguity, more conflicting information, more unhappy customers, more situations where the policy manual and the actual situation don’t line up – because AI didn’t just take the easy cases, it took the cases that were easy precisely because they didn’t require judgment. What’s left is, almost by definition, everything that does.
I’ve been thinking about this since finishing Migrating Scarcity. The book’s central claim is that technology doesn’t eliminate scarcity – it moves it. AI makes routine execution abundant, and scarcity migrates toward judgment. But there’s another consequence I didn’t spend much time on in the book. Scarcity can also move within the distribution of a single job – quietly making the average unit of remaining work harder, even as the total amount of work falls.
Which brings us back to her dashboard. Cost per contact fell because AI’s transactions are nearly free. Every blended metric that gets reported upward looks like a story about efficiency. But those averages increasingly describe the work the machine is doing, not the work her team is doing. An improving service metric doesn’t mean her agents got better at their jobs. It can mean the easy cases – the ones that used to pull the average down – aren’t reaching her agents at all anymore.
This is where it stops being a feeling she can’t name and starts being a decision someone else makes for her.
Her VP looks at the same dashboard and sees automation at 70%. The conclusion writes itself: if AI is handling seven in ten cases, the team needs roughly seven in ten fewer people. It’s a clean number, and clean numbers travel well in a budget meeting.
That conclusion might even be right, on headcount. A team handling 30 hard cases a day may genuinely need far fewer people than a team handling 100 mixed cases a day. The error isn’t necessarily the number. The error is the reasoning behind it – an assumption that capacity scales linearly with automation, while ignoring that the composition of the remaining work has changed. You can’t staff the residual 30% using the assumptions you used for the original 100%, because the residual 30% isn’t a smaller sample of the same job. It’s a harder job, done by fewer people, who now need more experience, more authority to deviate from the script, and more support than the average performer on an average day used to need.
Cut the staff by the automation rate and change nothing else, and you get a smaller team facing a harder caseload with the same training, the same authority limits, and the same thin margin for error the old team had – calibrated for a workload that no longer exists.
None of this will show up in the metrics that made the case for automation in the first place. Cost per contact will still look great. It will just be an increasingly accurate description of the transactions the machine is handling, and a decreasingly accurate one of the work left to the people still on the floor.
We keep asking what percentage of work AI will take from people. I’m starting to think that’s the wrong question.
If AI takes the easiest work first, the work left to people won’t simply be less of the old work. It will be different work.
And that means the average we’ve been managing may no longer describe the people we’re managing.
Vijay, your instinct is spot on (per usual). Economics teaches us that you should never pay attention to averages. You need to look at the margins. In your example, the service center was operating with a staff of 100 workers using their tools (supply curve) and had a specific caseload coming their way (demand). What’s important to understand is that introducing automation (AI or otherwise) only impacts the supply curve — not the demand curve (yet) and definitely not the quantity produced. So, by definition 70% automation will not necessarily reduce staff by 70%. It might, but that all depends on the impact the automation has on the supply curve.
How the supply curve is impacted by the automation is a function of the what was automated and how work is done in the factory. For example, let’s say you introduce automation into a unit of work that your team has always handled with a single individual. In others words, you don’t hand work from person to person to complete the job. Each widget in the “factory” is always made with one worker. There isn’t one big assembly line. There are 100 different assembly lines.
When you automate, you should expect a huge variability in the productivity across those 100 assembly lines. Some might be really fast at doing the tasks you took away, but really slow at the tasks that remain. Furthermore, you might introduce more quality variance because some of the people may be less skilled at the remaining tasks or exercise worse judgement, thereby increasing defects. So, if you run a factory like this, you should definitely not expect that 70% automation will result in 70% staff reduction. You’re dealing with 100 supply curves.
But, if the work performed by humans in the process you’re automating resembles more of an assembly line, the impact of automation will be entirely different. You have one supply curve you’re dealing with. And if you’re smart, you will already know which staff do the tasks that have been automated and you can be selective in removing them, while keeping the people on the assembly line who perform the non-automated tasks.
But, here’s the thing. You still shouldn’t expect a 70% staff reduction from a 70% automation. It will depend on how time intensive the automated tasks were. If each task was highly specialized (requiring different skills) but not very time consuming and most of the time for production happens at the non-automated nodes that remain, you will not see a 70% uplift in productivity. But, if the tasks you automated were highly time consuming and the remaining non-automated tasks are not, you will see a greater than 70% uplift.
My point is that you have to understand how work by the humans running the process you’re introducing agentic AI into. Of course, you could instrument the system and try to be precise about it, but I’m not sure that’s worth the effort. Instead, you could classify the work into either “craftsman-like” or “assembly line”. Once you do that you can at least directionally begin to understand the impact that automation may have.
But as a general rule, averages do lie to you! So, whatever you do, don’t trust them!
LikeLike
Agreed on the different work concept. The consequences of that different work is also different. It takes more effort, work fatigue sets in early and work satisfaction comes down. Basically the same feeling that comes in your life sometimes “Why Me?”… handling all corner cases all the time. As rightly mentioned we deal with different people with different set of needs and empowerment for which no one single dashboard exists to measure
LikeLike