← All insights

How to structure an AI pilot for professional services (without pilot theatre)


Most professional services firms do not have an AI adoption problem. They have a pilot-to-production problem. MIT’s Project NANDA found 95% of generative AI pilots show no measurable return on the P&L, and separate industry tracking shows the share of enterprises abandoning most of their AI programs jumped from 17% to 42% in a single year. The pilots themselves are rarely the issue - the demo usually works. What fails is the structure around it: no defined success criteria, no answer ready on data governance, and no plan for what happens if it works. If you’re weighing how to structure an AI pilot for professional services, that structure matters more than the tool you pick.

What “pilot theatre” looks like from the inside

Pilot theatre is a pilot that was never going to be evaluated against anything. A team runs a promising trial, people in the room like the demo, and three months later nobody can say whether it actually moved a number. There was no baseline to compare against, no owner accountable for the outcome, and no criteria written down before the trial started that would have forced a real decision - scale it, fix it, or stop. The pilot doesn’t fail so much as quietly dissolve, and the next one inherits the same scepticism.

For a partner-led firm, that scepticism is expensive. Every pilot that dissolves without a verdict makes the next funding conversation harder, because the mandate to try again gets weighed against the last one that went nowhere.

Define success criteria before you touch a tool

A pilot needs three things settled before it starts, not after:

  • A single business metric it’s meant to move - turnaround time on a document type, hours reclaimed per matter, error rate on a specific task. Not “efficiency.” A number you could measure last month and measure again in eight weeks.
  • A baseline, taken before the pilot starts, so “it feels faster” has something to be checked against.
  • A pre-agreed threshold for advancing - for example, a 15% improvement against baseline with no unresolved security or compliance flags. Write down in advance what result would justify moving to full deployment, and what result would justify stopping. Deciding that after the fact is how pilots drift into theatre.

Keep the pilot itself short - 60 to 90 days in one team or practice group, not a firm-wide rollout. A tightly scoped pilot with a clear metric is also a much easier approval to get in the first place than an open-ended technology investment.

Answer data governance before anyone asks

For a firm bound by APES 110 confidentiality obligations and client privilege, the governance questions are not a compliance afterthought to handle if the pilot succeeds - they’re gating criteria that need an answer before day one. At minimum: which system or model processes client data, whether it’s used to train anything beyond your own output, and where it’s stored and processed. From 10 December 2026, the Privacy Act’s automated decision-making transparency reforms require Australian businesses to disclose in their privacy policy when personal information feeds into automated decisions that affect people - so any pilot that touches client-facing decisions should already have that answer ready, not scrambled together after the OAIC’s final guidance lands.

If your team can’t answer those questions cleanly for the systems already in use, that’s the actual finding of the pilot - and it’s worth knowing before you scale anything, not after.

A structure that survives contact with the business

  1. Pick one workflow with a named owner - not a department, a person accountable for the result.
  2. Write the metric and the baseline down before any tool is switched on.
  3. Confirm what the pilot will touch - which systems, which client data, who else’s systems it reaches into.
  4. Set the threshold for advancing or stopping, agreed with whoever holds the budget decision.
  5. Time-box it, then hold the review against what was written down in step 2, not against how the demo felt.

That sequence is slower than picking a vendor and running with it. It is also the version that gives you an actual answer at the end, which is the entire point of running a pilot in the first place.

Where to start

The gating questions in step 3 - what data you hold, where it lives, and how clean it is - are usually the slowest part to answer honestly, because most firms haven’t looked. Our complimentary AI-Readiness Data Check is built to answer exactly that before you commit a pilot budget: twelve questions, about ten minutes, and a plain picture of your data and governance position going in. Treat it as pilot phase zero - the foundation check that makes the actual pilot’s success criteria mean something, rather than a substitute for running one.