All posts

What changed in model cost that makes this work now

The make-or-buy math on automating a document-heavy process flipped because the unit cost of running a model collapsed, not because any one capability arrived. Here is what a non-technical operator should take from the cost curve.

Samuel Mirpuri

Samuel Mirpuri

Co-founder & CEO of flowscope, previously leading digital transformations at McKinsey.

· Why now, and the operator's view

The same workflow that was uneconomic to put through a model in 2022 is economic now, and the reason is mostly arithmetic rather than any one headline capability. Stanford's Institute for Human-Centered AI, in its Artificial Intelligence Index Report 2025, tracked what it costs to run a model at a fixed level of performance and found that the price to reach the quality of GPT-3.5, held constant at a fixed score on a standard knowledge benchmark, fell more than 280-fold in just under two years. The figure went from around $20 per million tokens in November 2022 to around $0.07 per million tokens by October 2024. That is the number an operator should hold onto, because it is the number that changed the make-or-buy decision on a whole class of back-office work.

What a token is, in operator terms

A model does not process a document the way a person reads one. It breaks text into small pieces, called tokens, each roughly a few characters long, and the cost of running the model is charged by how many tokens go in and come out. A page of an invoice or a remittance advice runs a few hundred to a couple of thousand tokens. So when the report quotes a price "per million tokens," it is quoting the cost of processing on the order of a thousand pages of input. An operator does not need to track tokens directly. What matters is the direction and the size of the move. At $20 per million tokens, running a high volume of documents through a model every month was an expense large enough to wipe out the labor it was meant to save. At $0.07, the model cost becomes small enough to ignore against the wage of the person doing the rekeying.

Why a 280-fold drop changes the decision

A cost decline of a few percent does not move a make-or-buy line; it shaves a margin. A decline of more than two orders of magnitude moves the line, because it pushes whole categories of work from the wrong side of the calculation to the right side. Consider a process that reads many documents and rekeys their contents into a system: accounts payable matching an invoice to a purchase order, a collections clerk pulling figures off remittances, an onboarding analyst transcribing a vendor's tax and banking forms. The per-document model cost used to dominate that math, so the rational answer was to keep paying people. Once the per-document cost falls far enough, the wage becomes the dominant term again, and the rational answer flips. The capability to read an invoice did not appear overnight in late 2024. What appeared was the ability to run that capability across thousands of documents a month without the bill exceeding the salary line it offsets.

The same report sets out the supporting trends underneath the headline. Hardware costs fell about thirty percent per year, and energy efficiency improved about forty percent per year over the period it studied. Those are the inputs that compound into the inference price an operator actually pays. The Federal Reserve, in an April 2026 FEDS Note by Jeffrey S. Allen, Monitoring AI Adoption in the U.S. Economy, assembles several official surveys and finds the same broad climb in business use, which is the behavior to expect once the unit economics turn.

The adoption number that follows the cost number

When the price of running something falls this far, usage follows, and the data shows it did. The AI Index reports that seventy-eight percent of organizations used AI in at least one business function in 2024, up from fifty-five percent in 2023. A single year does not produce a fundamentally better model in every function at once, so this is not a story about a new feature arriving. It is a story about a threshold being crossed, where a cost that used to fail the business case now passes it, and processes that sat on a backlog for years became worth doing. For an operator wondering why the pressure to act has intensified, the cost curve is most of the answer, and the adoption jump is its consequence. We walk through how that backlog gets read and prioritized in the state of enterprise AI in 2026, and which specific tasks the falling cost actually puts in reach in what the agent-reliability curve says is automatable now.

The governance line moves with the cost line

The same report carries a second number that belongs next to the first. The AI Index recorded 233 AI-related incidents in 2024, the highest annual count it has tracked. The two findings are connected. When something becomes cheap enough to run at volume, it gets run at volume, including in places where the controls have not caught up. A model that reads invoices and writes the result into an accounting system is doing more than analysis. It is taking an action, and actions need the approval thresholds, logging, and exception handling that a person at a keyboard supplied implicitly and a script does not. The cost curve is what makes the work economic. The incident count is the reminder that economic and safe are different properties, and that the second question to ask, right after whether you can afford to run this, is what happens when it gets one wrong. We treat that as a design requirement rather than an afterthought in controls for autonomous actions, and the larger shift from systems that record to systems that act is laid out in the system of action.

A reasonable counter, answered

A reasonable counter is that the 280-fold figure is a benchmark artifact. It measures the cost of reaching a fixed, fairly modest performance level, and a document-heavy production workflow often needs a more capable, more expensive model than the cheapest one that clears the benchmark, so the real saving on a given job is smaller. There is genuine truth in this. The frontier of capability is not cheap, and the headline number flatters the math if you read it as the price of any task rather than the price of a held-constant one. But the operator's decision rarely sits at the frontier. The work that the cost curve newly makes economic is exactly the high-volume, well-bounded reading and rekeying that does not demand a frontier model, and for that work the standardized number is close to the right one. The expensive frontier matters for the hardest few percent of cases, which is a separate question, addressed in build, buy, or hire an AI capability. For the bulk of the document queue that has waited on a backlog for years, the line moved, and it moved enough to act on.