The clearest signal that enterprise AI buying is being rebuilt is the failure rate it is reacting to. MIT's NANDA initiative, in its State of AI in Business 2025 study, found that about ninety-five percent of enterprise generative-AI pilots delivered no measurable impact on the profit-and-loss statement, and the failures traced not to weak models but to integration and what the authors call the organizational learning gap. A procurement process designed to select against that rate cannot keep working the way it did for conventional software. The old request for proposal asked vendors to describe what their product can do. The new one asks them to prove what it did, on the buyer's data, before the contract scales. The rest of this explains the mechanism behind that change.
Why the feature-list RFP stopped predicting anything
A request for proposal built on advertised capability assumes the description and the result are close enough to substitute for each other. For deterministic software that assumption mostly held. If a system claimed to export to a given format, it either did or it did not, and a reference call settled the question. Generative systems break the substitution for two reasons. The same model that handles eighty percent of a workflow cleanly can fail the remaining tail in ways the demonstration never surfaces, and the part that matters most, integration into the buyer's existing systems and data, is exactly the part a vendor demo cannot show. The NANDA finding is the direct evidence: the pilots that failed were not selecting bad models, they were selecting on the wrong attribute. Gartner names the supply-side version of the same problem. In a June 2025 forecast it warned of agent-washing, the practice of relabeling existing software as agentic, and predicted that more than forty percent of agentic-AI projects will be canceled by the end of 2027. When the label is unreliable and the demonstration is unrepresentative, a procurement process that scores both is scoring noise.
What buyers are asking for instead
The move under way is from a feature checklist toward a bounded production pilot with predefined success metrics. Three changes follow from that. First, the evidence standard shifts from vendor benchmarks to reliability on the buyer's own data. A model's published accuracy on a public test set tells the buyer almost nothing about the documents, exceptions, and edge cases that live in their accounts-payable queue, so the credible vendor runs against a sample of the buyer's real workflow and reports the result. Second, the unit of evaluation becomes a real workflow run end to end rather than a capability shown in isolation, because the failures that sink projects appear at the seams between steps and systems, not inside any single step. Third, and operators tend to underweight this one, the operating model behind the software is weighted more heavily than the model inside it. The question is no longer only which model a vendor uses but who fixes the exception at 4 p.m. on close day, how the workflow gets corrected when it drifts, and whether the provider stays on the result or hands over a configuration and leaves.
What a defensible evaluation now contains
A modern evaluation has roughly four parts. It uses a real workflow, chosen because it is high-volume and currently manual, not because it demonstrates well. It runs against a held-out test, a slice of the buyer's own historical cases the vendor has not seen and has not tuned against, so the reported number measures generalization rather than memorization. It states an exception-rate target up front, the share of cases the system is allowed to route to a human, because no honest provider claims a workflow runs at one hundred percent autonomy, and the target is what makes the result falsifiable. And it carries a reversibility and audit requirement: every action the system takes must be logged, attributable, and undoable, so a wrong write can be traced and reversed rather than discovered in a quarterly reconciliation. The buyer who specifies these four can compare vendors on what they produced, which is the point of the exercise. We walk through how to construct one in how to tell a vendor that ships from one that demos, and the reason the exception-rate target matters so much sits in the 80-to-99% problem.
The governance frameworks already encode this
The audit and reversibility requirements are not flowscope inventions, and a buyer does not have to derive them from scratch. The NIST AI Risk Management Framework organizes the discipline around measuring and managing risk across a system's life, which presumes you can observe what the system does in production. ISO/IEC 42001, the management-system standard for AI, expects documented controls and ongoing oversight rather than a one-time sign-off. The EU AI Act's high-risk requirements go further, mandating logging, human oversight, and traceability for systems in regulated uses. A procurement process that asks for a held-out test, an exception target, and a full audit trail is encoding what these frameworks already require, ahead of the moment an auditor or a regulator asks for it. The specific controls that satisfy the reversibility half of this are the subject of controls for autonomous actions.
What this does to the vendor field
An empirical procurement process sorts vendors differently than a capability one does. A firm whose strength is the demonstration, the polished interface and the impressive benchmark, scores well on a feature-list RFP and poorly on a held-out test against messy real data. A firm that builds and operates running workflows scores the reverse. The shift therefore moves spending toward providers organized as services, the ones staffed to integrate into existing systems and to stay on the result, and away from the ones organized to license a product and exit. This is also why the internal build-versus-buy question is changing shape. The choice is less about owning the code and more about who can demonstrate a working result and operate it, which we treat in build, buy, or hire an AI capability.
A reasonable counter, answered
A reasonable counter is that bounded production pilots are slow and expensive, that running every candidate against a held-out slice of real data is a heavier procurement than most operators can staff, and that a feature-list RFP at least screens the field cheaply before the costly stage. There is something to this. A real pilot does cost more than reading a slide, and an organization that runs ten of them in parallel will spend more on procurement than one that runs none. But the comparison is not pilot cost against zero. It is pilot cost against the documented ninety-five-percent failure rate of the projects that skipped the pilot, each of which consumed integration time and internal credibility before it was canceled. Measured against that base rate, a bounded test on your own data is the cheap option, and it is the demonstration that turned out to be expensive.