Skip to main content
K4M2 AI

Practical guides

How to choose an AI consulting firm

Almost every selection process tests whether a firm can build. The thing that decides the outcome is whether they will tell you what not to build.

By Subhash Trivedi8 min read

Most AI selection processes are run like software procurement. A requirements document goes out, four or five firms respond, each demonstrates a working prototype, and the one with the most convincing demo wins.

That process tests the wrong thing. By 2026, technical capability is not what separates firms. Several can build a retrieval system, fine-tune a model, or stand up an agent. What separates them is whether the work reaches production, changes a number the business already reports, and keeps working after the engagement ends.

This is a guide to running that selection properly: what to establish before you brief anyone, what to ask, what should worry you, and what the first fortnight with a good partner looks like.

First, name the gap you are actually filling

The most common cause of a bad appointment is not a bad firm. It is a brief that does not say which of three different problems the firm is being hired to solve.

Judgement

You do not yet know where AI would pay. You need someone to find and size the opportunities, and to say which ones are not worth it. This is diagnostic work, and it should be short.

Capacity

You know what to build and cannot staff it. You need engineers who have shipped this class of system before, working inside your constraints rather than around them.

Operation

Something is already live and drifting. You need evaluation, monitoring, and the discipline to keep a system honest once real users are on it.

Firms are rarely equally good at all three, and the sales conversation will not volunteer which one they are weakest at. If you cannot name your gap, a short scoping exercise is a cheaper way to find out than a build. That is the entire purpose of a discovery engagement.

What the evidence says about outside help

MIT's The GenAI Divide: State of AI in Business 2025 is the study most often quoted for its headline finding, that around 95% of generative AI pilots deliver no measurable impact on the profit and loss statement. The more useful finding sits further down: tools built with an external partner succeeded roughly twice as often as internal-only builds.

That is not an argument that consultants are better engineers. It is an argument about forced specificity. An outside team has to establish, in writing, which workflow changes and who owns the result, before anyone is allowed to start. Internal teams are permitted to skip that step, because everyone assumes it is already understood.

McKinsey's State of AI survey and the UK government's AI adoption research point the same way from different angles: most organisations are still piloting, and readiness to scale lags adoption badly.

So the question to test is not can you build it. It is will you make us decide what we are building and why.

A firm that helps you skip that decision is selling you the 95% outcome with better slides. We wrote about the shape of the underlying problem in most companies don't have an AI strategy, they have activity.

Seven questions worth asking

Ask these in the room, and listen for whether the answer is specific or rehearsed. A good firm will have thought about all seven before you asked.

01

What have you told a client not to build, and why?

The single most revealing question. A firm that has never declined work either has no judgement or no leverage. Ask for the reasoning, not the anecdote.

02

Show me something you put into production and how you knew it worked.

Production, not pilot. The answer should include the measurement method and who at the client signed off on it, even if the client cannot be named.

03

Who exactly will be doing the work?

Names, seniority, and how much of their week. The classic failure is a senior team in the pitch and a junior team on delivery. Put the named people in the contract.

04

How do you choose a model, and what happens when a better one ships?

Listen for an evaluation method rather than a preferred vendor. If the architecture cannot survive swapping the model, you have bought a dependency, not a system.

05

What does it cost to run at real volume, not pilot volume?

Consumption pricing is cheap at pilot scale and rarely modelled beyond it. Ask for cost per decision at your expected volume and at three times it.

06

Who owns the code, the prompts, the evaluation sets, and the data?

Evaluation sets are the asset people forget to ask for. They encode what good looks like in your business, and they are expensive to rebuild.

07

What does handover look like, and when do you leave?

A firm that cannot describe its own exit is describing a retainer. Ask which of your people will be able to change the system without them.

Signals worth taking seriously

None of these is disqualifying on its own. Two or three together usually are.

01

An outcome guaranteed before any diagnostic work. Nobody can price a result they have not yet looked at.

02

A proposal that opens with a technology stack rather than a workflow. The order reveals what they think the problem is.

03

Case studies that describe what was implemented but not what changed. Implementation is an activity, not a result.

04

Reseller economics they will not disclose. If a firm earns margin on the platform it recommends, you need to know before the recommendation.

05

No answer on governance. Ask how they map to the NIST AI Risk Management Framework and, in Europe, to the EU AI Act.

06

Enthusiasm for everything. A firm with no opinion about where AI does not belong has not been doing this long.

Our own position on the last two is written down: what we will and will not take on, and how we handle responsible AI.

Large firm or small firm

The honest answer is that size is a proxy for two things you can measure directly: how much change management the engagement carries, and how close the people making decisions sit to the people writing code.

A very large programme touching many business units needs institutional weight. A single workflow that has to reach production this quarter usually does not, and the overhead becomes the cost. The MIT research noted that mid-market companies often reached full implementation faster, in around ninety days, because they could redesign a process without a committee.

Judge on the same evidence either way: named people, production references, an evaluation method, a stated exit. We wrote about the mid-market version of this in AI for mid-market enterprises.

What a good first fortnight produces

You do not have to wait six months to find out whether the appointment was right. Two weeks is enough, if you agree in advance what the two weeks must produce.

By the end of it you should be holding four things: one named workflow with its current cost or cycle time; the number that should move and the threshold that would justify scaling; the person inside your organisation accountable for that number; and the condition under which you would stop. If a firm cannot get you there in a fortnight, more time will not help.

Two related notes cover the measurement side of that: why AI ROI is an organisational problem, and the five places money leaks between pilot and return.

Choosing well is mostly a matter of refusing to be impressed. The demo will always work. Ask instead what the firm turned down, what they measured, who will actually be in the room, and when they plan to leave. The answers to those four are the whole decision.

Further reading

Forbes on MIT NANDA, The GenAI Divide: State of AI in Business 2025 ↗McKinsey, The State of AI ↗UK Government, AI adoption research ↗NIST, AI Risk Management Framework ↗European Commission, regulatory framework for AI ↗

More from us

Discovery: how we scope work before buildingHow we workMost companies don't have an AI strategy. They have activity.How to choose a model for production

Ask us the seven questions.

We would rather be assessed on them than on a demo. If you can name the workflow, we can tell you within a fortnight whether it is worth building against.