Blog/The third layer of consulting

How a forward-deployed engineering organization is actually built and staffed

The hard part of delivering working AI on a client's systems is staffing a role that combines production engineering, client elicitation, and judgment about exceptions, and that combination is what limits the model's scale.

The constraint on a software-delivered services model is finding the person who can build the software while sitting next to the people whose work it changes. A forward-deployed engineer has to do three things at once: write production code that runs against a client's actual systems, elicit from operators how the work happens rather than how it is documented, and decide which of the exceptions that come up in that work are worth a person's time and which can be handled by a rule. Each of those is a separate competence. The market for the first one is tight on its own, and the combination is rarer than any of its parts. That is the staffing problem, and it is the reason the model is hard to scale.

The three competences, and why they rarely coexist

Most engineering hiring optimizes for one axis, which is whether the person can ship reliable code. That is necessary here and not close to sufficient. A forward-deployed engineer who can write a clean integration but cannot sit with a controller and reconstruct how the month-end close actually runs will build the wrong thing accurately. One who is good in the room but cannot write a write-back into a system that has no usable API will produce a convincing demo that never reaches production. Both of those skills, present together, still leave the third gap, which is knowing, when a reconciliation will not tie or a document does not match its template, whether this is a case the agent should escalate or a case the agent should learn to clear on its own. That judgment is what separates a tool that handles the clean center of a process from one that holds up against the long tail of variation that real work produces.

The published accounts of the role agree on the embedding. Palantir's own descriptions of the forward-deployed engineer present it as embedding with a customer, understanding the domain problem, and writing the code to solve it, rather than handing the domain understanding to a separate delivery team. Joe Schmidt at Andreessen Horowitz, writing in June 2025 about why the forward-deployed engineer had become the hottest job in startups, stressed proximity: forward-deployed teams are closest to the customer, and the feedback loop between that front line and the product has to be tight. We have written before about why this is the third layer of consulting rather than a variant of the first two. The staffing consequence is what this post is about.

The labor market for the engineering half

Even the narrow engineering half of the role is expensive and scarce. Compensation data for 2025 places total compensation for senior machine-learning and AI engineers commonly in the $200,000 to $350,000 range and above. Levels.fyi, which aggregates self-reported offers, has AI-focused software engineers in the US earning $245,000 a year on average as of mid-2025, with the premium over non-AI engineers widening at each level of seniority, and the Stack Overflow Developer Survey 2025 has the specialization third among US developer salaries, behind only senior executives and engineering managers, with a median of $189,500. On the scarcity side, Bain & Company found demand for AI skills growing 21% a year since 2019, with compensation for those skills up 11% a year over the same period and a shortage it expects to persist through 2027. Treat the pay figures as aggregator ranges rather than a fixed number, because they vary by location and by company tier, but the direction is not in dispute. You are competing for this talent against firms that will pay at the top of that band for the engineering skill alone.

Now add the requirement that the same person be willing to spend days inside a client's office watching accounts-payable clerks work, and be good at it. That second requirement narrows the pool again without lowering the price, because many strong engineers neither want client-facing work nor are good at it, and many strong client-facing people cannot write production code. The combination is the binding constraint, and no amount of model capability relaxes it.

How the team is structured around the combination

There are two ways to put client-facing judgment and engineering depth in the same place. The first is to hire individuals who carry both, which is the purest form of the role and the hardest to staff at volume. The second is to build small pods of two or three people, where the competences are adjacent rather than resident in one person, and the pod as a unit owns an engagement end to end. The pod form is more tractable to hire for, but it only works if the people sit close enough that the domain understanding does not get lost in a handoff. The failure mode of large delivery organizations is precisely the handoff, where the people who understood the workflow write it down and the people who build work from the document. We have written about why documented processes rot; a structure that depends on a written handoff inherits that decay.

The structural test, whether the unit is a person or a pod, is whether the same locus of accountability spans elicitation through production. If it does, the judgment about which exceptions matter is made by someone who has seen both the workflow and the code. If it does not, that judgment gets made twice, badly, by two groups who each see half.

Internal tooling is what lets a small team carry many engagements

A scarce, expensive team can only be economic if each member carries more engagements than a traditional consultant does. That throughput comes from internal tooling: the capture stack that records how work happens, the harness that turns observed sequences into a candidate agent, and the deployment and monitoring rails that let one engineer watch several agents in production without re-deriving the plumbing each time. The tooling replaces the repetitive engineering around the judgment, not the judgment itself, so that the rare person spends their time on the part only they can do. This is the mechanism behind why services-as-software firms scale at all: by raising the number of engagements each scarce person can hold, rather than by hiring proportionally to revenue.

There is a second, less obvious effect. Brynjolfsson, Li, and Raymond, in "Generative AI at Work" (QJE 2025), found that an AI assistant trained on the behavior of the most effective workers raised the productivity of everyone else by propagating those workers' practices. The same logic applies inside the delivery organization. The judgment of the best forward-deployed engineers, about which exceptions to escalate and how to structure a write-back, can be encoded into the tooling so that a newer engineer inherits it rather than rediscovering it. That is how a small expert team extends its judgment beyond the hours its experts personally work.

A reasonable counter, answered

A reasonable counter is that this combination is not actually rare, that plenty of engineers can talk to customers, and that the talent constraint is a story firms tell to justify high fees. There is something to it. Many engineers are perfectly capable in a room, and the elicitation skill is teachable to a degree. But the test is harder than attending a client meeting: sitting through a process, noticing the step the operator does not mention because it is automatic to them, and deciding on the spot which deviations the agent must handle. Firms that staff the engineering and the client work separately keep proving the point, because their deliverable drifts toward analysis when the people who understand the work are not the people who build, which is the same pattern that determines the operator's role after automation. The technology is buyable. The organization that runs it on someone else's systems is the part that has to be built, and it does not come off a shelf.

Book a demo

Start with one department

We map how one team’s work runs, show you where the hours go, and automate from there.