All posts

Why an agent does discovery and delivery better than a consultant does either

A consulting engagement loses information at every handoff between discovery, redesign, build, deployment, and monitoring. One agent holding the whole arc removes the self-report gap and the handoffs at once, which is where the accuracy and the speed come from.

Samuel Mirpuri

Samuel Mirpuri

Co-founder & CEO of flowscope, previously leading digital transformations at McKinsey.

· The value shift: code to context

A consulting engagement that maps a process, redesigns it, builds the automation, deploys it, and then monitors it in production is five distinct activities, usually performed by different people, often in different firms or different time zones. Each activity reads the output of the one before, re-derives some of what that output left implicit, and writes its own document for the next person to read. Flowscope runs the same arc as one agent that observes the work, drafts the redesign, ships the write-back, and watches it in production. The case for doing it that way rests on two specific losses that a chain of human specialists incurs by construction: the gap between what people say they do and what they actually do, and the information that leaks at every handoff between phases. Neither loss has anything to do with whether an agent writes better code than an engineer.

What the interview gap actually costs

A traditional discovery phase relies on self-report. A consultant interviews the people who run a process and writes down what they describe. The difficulty is that practiced work is largely tacit. Michael Polanyi's starting point in The Tacit Dimension, that "we can know more than we can tell," is the interviewer's problem in one line: people skip the steps they have automated in their own heads, omit the exceptions they handle without thinking, and describe the official procedure rather than the one they use when the official procedure is too slow. The map that results records what people believe they do, filtered through what they choose to tell an outsider. The belief itself is a reconstruction. Richard Nisbett and Timothy Wilson showed in a 1977 paper in Psychological Review that people reporting on their own mental processes draw on plausible theories of what they must have done rather than on any direct introspection.

Observation removes that filter. An agent that watches the actual sequence of actions, the applications touched, the fields edited, and the order in which they happen records behavior rather than testimony. Flowscope made this case for discovery specifically in Discovery doesn't need three months: the observed trace is the map, and it is more accurate than the interviewed one because there is no self-report step in which information is lost. The same advantage extends past discovery, because every later phase inherits whatever the map got wrong.

Two consultants, two different maps

Even setting aside the interview gap, the human chain has a second accuracy problem, which is that authored artifacts vary by author. Give two competent consultants the same set of interviews and you get two different process maps, because a map is a model and modeling involves judgment about what to include, how to name a step, and where one activity ends and the next begins. The variance shows up even under controlled conditions: a study in the journal Information Systems gave 89 novice analysts one identical written scenario, an airport check-in and boarding process, and their diagrams split into five distinct design archetypes, from textual descriptions through hybrid forms to graphics. Training standardizes the notation but does not remove the judgment. That variance does not average out across the engagement. It is a structural feature of having a person translate observation into a document, and it compounds, because the redesign is built on the map and the build is built on the redesign.

An agent that produces the map from the trace produces it the same way every time, against the same definitions. The redesign it drafts references the specific observed steps rather than a re-described summary of them. When the build writes back into the customer's systems, it writes against the same representation the map and redesign were derived from. So the map, the redesign, and the write-back are three views of one underlying object rather than three documents that a reader has to reconcile.

Where the compression comes from

The speed advantage comes less from the model being fast than from the elimination of handoffs. In the consulting chain, discovery hands a deck to the redesign team, the redesign hands a specification to the build team, the build is frequently handed to an offshore engineering pod, and the deployed system is handed to whoever monitors it. Each handoff is a point where the receiving party reads a partial description of what the previous party knew and re-derives the rest. Time is lost in the re-derivation, and accuracy is lost because the re-derivation is a guess at the original intent. The deck-as-deliverable problem flowscope described in Your AI consultants left you a deck is the same problem viewed from the customer's side, where the deck is the residue of a handoff that was supposed to carry the knowledge and did not.

One agent holding discovery through delivery has no handoffs to lose information across. It does not write a deck for a redesign team because it is the redesign step, and it does not write a specification for a build pod because it is the build step. The compression comes from removing the boundaries between phases rather than from any single phase running faster, since the boundaries are where most of the time and most of the error accumulated. This is the broader case made in Forward-deployed agents, not just engineers: the unit that ships is the agent that did the observing, not a separate team that received its notes.

The decay between phases

There is a third loss that the continuous artifact removes, which is decay over calendar time. A consulting arc that runs for months allows the underlying process to change while the engagement is still running. The map is accurate as of the interviews, the redesign is built against a map that is now weeks old, and by the time the build deploys, the process it was designed for has drifted. An agent that observes, redesigns, and deploys on a compressed timeline reduces the window in which drift can accumulate, and because it remains in production monitoring the running system, it observes the next drift as it happens rather than waiting for the next engagement to discover it.

What stays human

None of this argues that the agent should decide what the redesigned process ought to be, or that it should be accountable for the outcome. The redesign decision is a business judgment about which steps to keep, which to cut, and what risk is acceptable, and that judgment stays human. The accountability for what the deployed system does stays human as well, because production is where the reliability problem lives. Flowscope made this argument in The 80-to-99% problem: getting an automation from a demo that mostly works to a system that works reliably enough to leave running is the genuinely hard part, and it is why the agent has to be bounded, with the high-variance exceptions routed to a person rather than guessed at. The continuous artifact makes the bounding cleaner, because the same representation that produced the map also defines where the agent's competence ends and a human's begins.

The case for specialists

The strongest objection is that specialization exists for a reason, since a dedicated redesign strategist or a senior build engineer is better at one slice of the arc than a generalist agent is across the whole of it, and the human chain trades handoff loss for depth at each step. That trade holds where the slice genuinely requires scarce human judgment, which is exactly the redesign decision that stays human. For the mechanical work of the arc, the observing and mapping and drafting and writing back and monitoring, the depth a specialist adds is smaller than the information a handoff destroys, because that work is execution against a known representation rather than novel judgment. The agent wins on the parts that are execution and defers on the parts that are judgment, and the engagement is structured so the human spends time on the decision rather than on re-deriving what the previous phase already knew.