All posts

Why your AI consultants left you a deck

By month six of a typical Big-4 AI engagement, the deliverable is a deck and a pilot proposal, with nothing written to a system of record. The model serves the customer it was built for. Most operators are not that customer.

Samuel Mirpuri

Samuel Mirpuri

Co-founder & CEO of flowscope, previously leading digital transformations at McKinsey.

· The AI consulting critique

The arc of a Big-4 AI engagement is familiar enough to narrate in advance. The firm wins the AI work, a six-month engagement is scoped, and the opening conversation is hopeful. By month three, executive sponsors are reviewing a slide deck of opportunities; by month six, a final deck of recommendations and a pilot proposal. At no point in the engagement did anything write to a system of record. The customer is left with strategic clarity, a portfolio of next-phase initiatives, and the option to either commit another half-million dollars or quietly let the program lapse.

Calling this a failure of execution misreads it. The consultants on these engagements are smart, the methodology is rigorous, and the analysis is genuinely useful for the work it was built to do. The deck is good work, and the deck is the deliverable by design.

Who is in the room on a Big-4 AI engagement

To see why, walk through the org chart on a typical AI engagement at one of the named firms. A partner owns the relationship, a principal runs the engagement, and an engagement manager keeps the workstreams aligned. Four to six consultants, some of whom carry the title "AI specialist," do the analysis, and behind them sits an offshore delivery center that nominally produces working artifacts. Everyone in that chart is billable, and everyone is on a six-month clock.

What is missing from the engagement team

There is no engineer in the customer's environment with permission to write to the customer's systems. The on-site team is staffed by management consultants whose toolkit is interviews, frameworks, and presentation. The offshore team can produce code, but they do it from a delivery center with limited environment access and a separate priority queue, often working from a different time zone with handoffs every twelve hours. The step from "we know what the agent should do" to "the agent is running in production on the customer's stack" is owned by no one in the room.

The six-month sequence: a deck at every phase

The six months divide into five phases. Discovery, the first, runs stakeholder interviews and current-state process mapping and identifies the opportunities; prioritization then sizes those opportunities into business cases and a portfolio view. The design phase lays out a future-state target operating model, a technology selection, and an implementation roadmap, and pilot definition follows with scope, success criteria, and a governance structure. Every one of those phases delivers its output as a deck. The fifth phase, if it survives the budget review, is the pilot itself, built by the offshore team against a thin-sliced version of the problem, and it closes with a demo rather than a deployment.

What is in production on day 181, in the great majority of cases, is nothing. Recommendations have been accepted, governance committees have been stood up, and budget has been allocated for a future phase, but no system of record behaves any differently than it did at kickoff. McKinsey's own published number is that roughly seventy percent of transformations fail to meet their objectives, and the Big-4 firms know it well enough to cite it themselves in pitches for the next round.

Why the model produces decks instead of software

Why the model produces this outcome is not mysterious. The firms that deliver AI engagements at this scale are professional services firms whose P&L is built on billable hours of senior consultants doing analysis rather than engineering. When the engagement requires engineering, it gets routed to a delivery center where the unit economics, the access, and the institutional priority are all different, and the two halves of the work stay structurally separate inside the firm because they are commercially separate. The senior consultants get rewarded for selling the next phase, the offshore engineers for landing the code the consultants specified, and nobody in the structure for ensuring that the agent is running in the customer's production environment six months from now.

How the Big-4 firms are restructuring

The Big-4 firms are aware of this and are restructuring under the pressure. McKinsey now reports 25 percent of projects priced against outcomes rather than time, and Fortune reports the firm's headcount fell by more than 10 percent in eighteen months, from about 45,100 at the end of 2023 to roughly 40,000, a decline the firm attributes to normal attrition and performance-review firings. Bloomberg reported in December 2025 that McKinsey's leadership has discussed cutting about 10 percent of headcount in non-client-facing departments, which could amount to a few thousand roles. Lilli, the firm's internal AI platform, is in monthly use by over 75 percent of employees. EY has hired 61,000 technologists since 2023, fifteen percent of the workforce, and is openly exploring what it calls service-as-a-software. PwC cut graduate hiring by thirty percent over three years. The firms are not blind to the model breaking; they are restructuring around it as fast as a professional-services P&L built on billable headcount can be restructured, which is to say, slowly.

The customer who needs AI in production this quarter cannot wait for the firms to finish restructuring. The structural reason the deliverable is a deck is that the people in the room write decks, and after fifteen years of training they are very good at it. Asking them to ship software is asking the wrong people to do a job their commercial structure was never designed to deliver, and there is no individual fault in that; the fault sits in expecting a model built for one job to do a different one.

Software, for its part, requires a different workflow, a different access pattern, and a person with deploy rights sitting next to the system administrator at the moment of deployment. None of that fits inside a Big-4 engagement model, and that is fine, because Big-4 engagements were not designed to do it. The model serves a board that wants strategy, an executive team that wants alignment, and a planning function that wants an opportunity portfolio; these are real needs, and the model serves them well. What the model does not serve is a CFO who needs the AP queue processed with one fewer person on the team this quarter.

Complaining about the consultants misses the point; the model serves the customer it was built for. The useful response is to recognize that AI delivery to operating businesses is a different job, one that needs its own delivery model, and to stop being surprised when the existing models do not deliver it. The foundational case for that kind of delivery model was made in 1990, by a man named Michael Hammer, in the pages of Harvard Business Review, and his argument about automation is where the redesign work has to start.