The shift: AI breaks consulting's production economics
Every consulting firm in the world runs on the same cost structure: every hour of output requires a person. An analyst costs $80,000 to $150,000 a year plus overhead whether they’re billable or not, so the whole industry is managed around utilisation, and growth always means hiring. The pyramid exists to arbitrage the gap between what a junior costs and what their work bills out at. David Maister formalised it decades ago: profit per partner is margin times productivity times leverage, where leverage is the ratio of juniors to seniors.
AI breaks the first premise. Output no longer requires a person.
What the production layer actually is
The work that consumes most junior and mid-level consultant time is structured production: decomposing a problem into testable branches, executing the analysis (the financial model, the market sizing, the competitor scan), and synthesising the result into a deliverable. In digital consulting the artefacts are even more mechanical: requirements documents, technology roadmaps, business cases, architecture options papers, vendor evaluations, governance packs, status reports. All of it is analytical production in a structured language, and all of it is exactly what AI agents now do well.
What AI cannot do is know whether the answer is right. A market sizing can pass every internal logic check and still rest on the wrong framing of what the client actually needs to decide. Code has automated tests; a strategy recommendation has no equivalent, so the validation mechanism for consulting output is senior judgment, and it can’t be automated away without removing the grounding for everything underneath it. That’s the durable human layer, and it’s what this venture staffs.
What the cost change looks like
The numbers are rough but the direction isn’t subtle. A junior analyst doing 40 hours of research costs a client roughly $6,000 to $12,000 through a traditional firm; the equivalent AI production run costs tens of dollars in compute. A full strategy engagement that a team of six delivers over three weeks at $250,000 to $500,000 has an AI production cost in the hundreds. Two orders of magnitude, sometimes three, and inference prices for a fixed level of capability have kept falling every year since 2022.
Senior time and the engineering effort to embed AI into a client’s real operations still cost real money, which is why the financial model in section 8 carries them honestly. But the production layer itself, the thing pyramids were built to staff, is now close to free.
What near-zero production cost changes
Utilisation stops mattering because there’s no standing army to keep busy; you pay for work consumed, not capacity waiting. Running ten investigative tracks in parallel becomes a decision about quality rather than budget, which matters because consulting problems are open-ended and the right framing is usually found by investigating widely, not by guessing well up front. Scaling decouples from hiring: more clients means more compute and a slightly fuller senior calendar, not an 18-month recruit-and-train lag. And pricing power moves to value. A cost base five to ten times lower first lets us serve clients profitably that the incumbents can’t reach at all — the big firms are simply too expensive for them, and the mid-tier, still heavily people-based, mostly aren’t built to — and then, where we do overlap, lets us price to the value delivered rather than an hourly rate while still out-margining a pyramid.
Why incumbents won’t simply do this too
The deepest reason isn’t that they can’t act — it’s that most of them have mis-sized what’s happening. The prevailing view inside the big firms is that AI is a five-to-ten-percent productivity gain for their consultants: a faster first draft, a quicker deck, the same work done a little sooner. We think that’s wrong by an order of magnitude. This isn’t ten percent off the cost of a deck — it’s one person doing the work of ten, and the production layer ceasing to be a staffing problem at all. A firm that treats a step-change as an efficiency tweak is optimising the wrong thing entirely.
And the firms that do grasp the magnitude are trapped anyway. Responding properly means dismantling the pyramid that pays their current partners, so the rational short-term move is to layer AI onto the existing structure and protect the revenue engine. That is what’s happening: the major firms are cutting graduate intakes (reported reductions of roughly 30% at PwC and KPMG, 18% at Deloitte, 11% at EY) while leaving their economics intact. Cutting the junior layer without changing the model shrinks the pyramid; it doesn’t replace it.
The window this opens has a clock on it. Industry analysis puts an 18-month horizon on providers demonstrating genuine AI-native delivery before they’re excluded from new mandates, with full market repricing of consulting fees playing out over several years. New entrants who are AI-native from the first engagement carry none of the structure that stops incumbents from acting, and the first credible players in each niche will set the reference point clients use to judge everyone else.
← Executive summary The model: advisory-native AI, built to operate →