When an automation goes into production, the work that determines its cost is not the work it does well. A system that posts an invoice, validates an expense, or runs a pay period handles the bulk of its volume without a person involved, and that clean path is close to free at the margin. The cost lives in the cases the system cannot finish on its own and routes to a human. Ardent Partners, in its 'State of ePayables 2023' report, puts the average invoice exception rate at 20.7 percent. The GBTA Foundation, working with HRS, found that about nineteen percent of expense reports contain errors that need correcting. Ernst & Young, in work commissioned by Paycom, found employers make an average of fifteen payroll corrections per pay period. One in five, roughly, is the number to keep in view, because it is the number that governs the economics.
The cost is in the tail, not the throughput
Start with what the clean path actually costs. A model reading an invoice, matching it to a purchase order and a receipt, and writing it back to the ledger consumes a fixed amount of compute per document and nothing else. Run a hundred clean invoices or ten thousand, and the per-document cost barely moves. That is the property that makes automation attractive, and it is also why the clean path is the wrong thing to optimize, since there is little left to take out of it.
The tail behaves differently. Every case that the system routes to a person carries a human cost to resolve, and that cost does not fall as volume rises. The Medical Group Management Association puts the average cost to rework a denied healthcare claim at $25.20. The GBTA study attaches roughly $52 to correcting one expense report. The Ernst & Young figure for payroll is $291 per correction. These are the unit costs of the tail, and because they are paid per exception rather than amortized across throughput, they are what a finished automation is really buying down.
The mechanism, written out
The marginal cost of an automated process can be stated directly. It is approximately the exception rate, multiplied by the cost to resolve one exception, divided by the degree to which a person's judgment carries across many cases rather than one. Call the last term the leverage of a human decision. A reconciliation analyst who resolves one mismatch and moves to the next has leverage of one. The same analyst who writes a rule that resolves that class of mismatch for every future occurrence has leverage well above one, and the divisor in the cost expression grows accordingly.
Two of the three inputs are fixed by the work itself. The exception rate comes from the process and the model; the cost per exception comes from what resolving one actually takes. The third input, leverage, is the one a designer controls, and it is where most of the difference between an economic automation and an uneconomic one is decided.
Why the exception rate dominates
The exception rate enters the cost expression as a multiplier on the entire human side of the work, which is why moving it has more effect than speeding up the clean path. Take the accounts-payable figure. Cut the 20.7 percent exception rate to ten percent and you have roughly halved the volume reaching the expensive part of the system, and you have done it without touching the cost per exception or the leverage term. Halving the time the clean path takes, by contrast, reduces a cost that was already small. The reliability of the model is what sets the exception rate in the first place. A more reliable extractor or matcher sends a smaller fraction to the tail, which is the practical reason model reliability shows up as a line on the bill rather than a line on a spec sheet. We have walked through how that reliability gets built in how production reliability gets engineered, and why only part of a process clears the bar for automation at all in what the reliability curve says is automatable.
When the queue costs more than the manual process
A poorly designed exception path can make the whole automation cost more than the manual process it replaced, and the mechanism is worth being precise about. The danger is that the cases routed to a person are systematically the hard ones. The clean path absorbs the easy invoices, the well-formed expense reports, the pay periods with no changes, and what reaches the human queue is the residue: the malformed, the ambiguous, the genuinely exceptional. If the cost to resolve one of those is higher than the average cost of handling a case manually, and if the queue strips away the context that made manual handling efficient, then the per-exception cost rises even as the exception rate falls. Multiply a smaller rate by a larger unit cost and the product can exceed where you started. This is the trap behind a familiar result, an automation that handles eighty percent of volume and still fails to pay for itself, because the twenty percent it could not handle got more expensive to resolve rather than less. The closure of that gap is the subject of the 80-to-99% problem.
Designing for leverage, not for volume
The design objective that follows is to make each human decision resolve many future cases rather than one. When an analyst handles an exception, the system should capture what made it an exception and what resolved it, so the next instance of that pattern clears automatically. A mismatch caused by a vendor that ships partial orders is not one exception. It is a recurring class, and the first person to resolve it should be the last. This is why the structure of the exception queue matters more than its length. A queue that presents each case in isolation has leverage of one by construction, and a queue that lets a person teach the system raises the divisor in the cost expression with every decision.
There is direct evidence that handing people the routine share raises their effectiveness on what remains. Brynjolfsson, Li, and Raymond, in Generative AI at Work, found that customer-support agents with access to an AI assistant resolved more cases per hour, with the largest gains among the least experienced workers, in part because the assistant encoded the practices of the best performers and spread them. The same dynamic governs an exception tail. A person freed from the clean path, and equipped with a system that learns from each decision, resolves more of the residue and resolves it better. We worked a concrete instance of this in three-way match, rebuilt, where the exceptions that survive matching are exactly the cases where judgment earns its keep.
A reasonable counter, answered
A reasonable counter is that this framing overstates how compressible the exception rate is. Some exceptions are irreducible. A customer who pays the wrong invoice, a tax jurisdiction that changed its rule mid-period, a genuinely novel document: no model drives those to zero, and pretending otherwise sells a number that will not be hit. There is real truth in that, and a designer who promises a one-percent exception rate on a process whose irreducible floor is eight percent is setting up a miss. But the published rates are not floors. An exception rate of 20.7 percent in accounts payable, or nineteen percent on expense reports, is not the irreducible residue of hard cases. It is mostly the easy-to-mechanize, a missing field, a tolerance that needs widening, a vendor record that needs cleaning. The irreducible tail is real and should stay with people. The work of the redesign is to make sure that is the only thing left in the queue, and that each decision a person makes there resolves more than the case in front of them.