All posts

What the operator's role becomes after the manual work is automated

When an assistant compresses the experience curve, the residual human work moves toward judgment, exceptions, and oversight. The labor economics says this is measurable, and it tells you who to hire and how to train.

Samuel Mirpuri

Samuel Mirpuri

Co-founder & CEO of flowscope, previously leading digital transformations at McKinsey.

· Why now, and the operator's view

The clearest evidence on what happens to a team after software takes over its routine work comes from a customer-support function, a setting most operators would recognize. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond studied support agents at a software firm who got staggered access to a generative-AI assistant, and they reported the results in 'Generative AI at Work' in the Quarterly Journal of Economics. The average worker resolved fifteen percent more issues per hour once the assistant was live. That headline number is the least interesting figure in the paper. The distribution underneath it is what should change how a CFO or COO thinks about hiring, training, and what the remaining people on the team are actually for.

The gain is not evenly distributed

The fifteen-percent average is the net of two very different effects pulling in opposite directions. Novice and lower-skilled agents resolved about thirty percent more issues per hour. The most experienced and highest-skilled agents barely moved, and at the very top the authors found small declines in measured quality. The assistant did not lift everyone by a fixed amount. It compressed the gap between the bottom and the top of the skill distribution, raising new workers toward the level the veterans had already reached.

The mechanism Brynjolfsson and his coauthors identified explains why. The model had been trained on the records of past conversations, including those of the firm's most effective agents, so the suggestions it produced encoded how the best workers handled a given situation. A new agent reading those suggestions was receiving the practices of people who had spent years learning the job, at the moment of the call. The authors connected this to a long-standing idea in economics about tacit knowledge, the kind of skill Michael Polanyi described in 'The Tacit Dimension' as knowing more than we can tell, and to David Autor's work on how tasks rather than whole jobs are what get automated or augmented. The hard-won, hard-to-articulate knowledge of an experienced operator had been made partly legible and partly transferable. That is the finding to build an operating model around.

What is left for a person to do

If the assistant raises a new worker toward the top of the experience curve on precedented cases, the question is no longer how to get a team to handle volume. It is what the team does with the cases the assistant cannot. Brynjolfsson and his coauthors found the assistant helped most on moderately rare problems, the ones an individual agent sees too seldom to have learned but that appear often enough across the firm's history to be well represented in the training data. The most common problems saw the smallest gains; the authors' explanation is that agents are already capable of addressing the problems they encounter most frequently. And where a topic had too little training data, the system might offer no suggestions at all. Those are the cases where no clean precedent exists: a customer with a contractual edge case, a problem that spans two systems neither of which holds the whole picture, a situation where the right answer turns on a judgment about intent or risk that no prior transcript settles.

This is the same pattern wherever routine work gets mechanized. The agent or assistant absorbs the well-precedented middle of the distribution, and what remains for a person concentrates at the exceptions and in the supervision of the system itself. In an accounts-payable function it is the invoice that will not match; in a collections process it is the account where the dispute is real rather than a stall. The residual human work is denser in judgment per hour than the work it replaced, because the easy cases have been removed from the queue. An operator who plans for this designs the remaining roles around exception-handling and oversight, not around throughput.

The deskilling risk is real, and it has a cause

There is a version of this that goes wrong, and the same body of research names it. Shakked Noy and Whitney Zhang, in their experiment published in Science, gave writing tasks to professionals with and without ChatGPT and found that the tool cut the time to complete a task by roughly forty percent while raising the average quality of the output. The catch is in how the gain was produced. The tool substituted for the worker's own effort rather than complementing the worker's skill. People wrote less and edited more, and on average the work got faster and better, but the worker was doing less of the thinking. Where Brynjolfsson and coauthors found the assistant propagating expert practice to novices, Noy and Zhang found a setting in which the assistant did the task instead of the person.

The difference between those two outcomes is not the model. It is the workflow around it. When the system surfaces a suggestion that the worker must evaluate, accept, or override against a case the worker understands, the worker keeps building judgment and the tool augments skill. When the system produces a finished output that the worker approves without examining it, the worker's own capability atrophies and the tool substitutes for effort. The small quality declines among the most experienced agents in the QJE study are an early sign of this: a veteran who defers to a suggestion that is good enough on average can do slightly worse than the veteran working unaided on the cases where the average is wrong. Designing for augmentation rather than substitution is a choice made in how the work is structured, and it is the central choice in any redesign worth the name.

Designing the roles that remain

The practical consequence for hiring is that the entry-level job changes character rather than disappearing. The thirty percent novice gain means a newer person becomes productive faster, which lowers the cost of bringing someone in, but it raises the importance of what they do once the routine cases are handled for them. The roles to design for are the ones the system cannot fill. The first is the person who owns the exceptions and has the authority to resolve them; the second is the person who supervises the system's outputs and catches the cases where its average judgment is the wrong judgment. Both are oversight roles in the sense that matters. One oversees the cases that fall outside the precedent, the other oversees the system itself.

Training has to follow. If the assistant has absorbed the tacit knowledge that used to live only in the veterans, then preserving that knowledge in people now requires deliberate effort, because the cases that built it no longer cross a junior person's desk. The dynamic that makes a new worker productive faster also removes the repetitions that used to teach them. A team that wants its people to retain the judgment the exceptions demand has to route hard cases to humans on purpose, which is partly why some processes should not be fully automated even when they could be, and why the documentation of how work is really done matters more, not less, after deployment.

A reasonable counter, answered

A reasonable counter is that the support-floor result will not generalize, because customer support is unusually well-suited to this kind of assistance, with high volume, repetitive cases, and a rich history of past conversations to train on. Plenty of operating work is messier than that, with thinner precedent and higher stakes per case. There is real force to this. The study's gains depend on the model having adequate training data, so a function with a thinner history of past cases gives an assistant less to learn from.

But the counter cuts toward the argument rather than against it. The thinner the precedent and the higher the stakes, the more of the residual work is exception-handling and oversight, which is precisely the work the evidence says stays with people. In a high-precedent function the human role narrows to a smaller set of genuine exceptions; in a low-precedent one it stays broad, because more of the cases are exceptions. The composition shifts in the same direction in both, toward judgment and supervision, and the operator's task is to staff for that shift rather than for the volume the system now absorbs. The mechanism that makes this work, and the aligned engagement that pays for outcomes rather than hours of analysis, both rest on watching how the work is actually done before deciding which parts a person should keep, which is why the redesign starts with shadowing rather than surveillance.