AI 20/20 Subscribe

FIELD GUIDE · AI VENDORS

How to evaluate an AI vendor

You are being sold to by people who cannot show you production. This is the question list for the next pitch meeting: what to ask, what a real answer sounds like, what a dodge sounds like, and where these deals actually go wrong. All of it is on this page. No gate, no email wall.

COMPILED 2026-08 · SOURCES LINKED IN PLACE · 10 QUESTIONS, 6 FAILURE MODES

Every week another vendor demo lands in your inbox, and the demo is impressive. The demo is always impressive. The question is whether anything behind it survives contact with your data, your volumes, and your compliance team, and the pitch is built so you find out after the contract is signed rather than before.

MIT's 2025 State of AI in Business review put a number on the gap: across some 300 public enterprise deployments, 95 percent of generative AI pilots produced no measurable P&L impact. The pilots did not fail because the models are weak. They failed because in the meeting, the buyer could not tell which vendor had done this before and which was reading from the same demo script as everyone else.

The questions below separate them. Ask them in order, and write the answers down. A vendor with a real product will enjoy this conversation. That reaction is itself a signal.

The ten questions

ASK IN ORDER

Q.01

Is this running in production anywhere today?

Not piloting. Not rolling out. Running, with real users, on real data, for months.

A real answer

Names a deployment and walks it end to end: how long it has run, at what volume, what it replaced, and what the on-call rotation looks like when it breaks.

A dodge

Pilots, waitlists, and design partners who cannot be named. A pilot is not production. Production has a pager.

Q.02

What does it cost to run, all in?

The license is the visible part. The rest is where budgets double.

A real answer

A number with parts: seats or license, usage or inference, the integration work, and the human review time your side carries after rollout.

A dodge

A per-seat price with everything else answered as "it depends." It does depend. That is why you are asking now, in the meeting, and not at renewal.

Q.03

What happens when it is wrong?

Every AI system is wrong some of the time. The vendor's relationship with that fact is the product.

A real answer

An error rate they measure, a definition of what counts as wrong, an escalation path, and a straight answer on liability.

A dodge

"The model is very accurate." In February 2024 a Canadian tribunal held Air Canada liable for a refund policy its own chatbot invented. Your output, your liability. Ask who carries it.

Q.04

What data do you need from us, and where does it go?

This question ends more deals than price does, which is why it comes fourth and not tenth.

A real answer

Named systems, retention windows, tenancy, and a plain yes or no on whether your data trains anything, with a data-flow diagram they already had ready.

A dodge

"We take security very seriously," followed by no diagram. Seriousness has paperwork.

Q.05

Which model is underneath, and what happens when it changes?

Model upgrades change behavior in ways no demo shows. The vendors who know this have scars, and evals.

A real answer

Pinned versions, an evaluation suite that runs before any upgrade ships, and a regression story they will tell you unprompted. Morgan Stanley gated its advisor assistant on evals before rollout. That is what care looks like.

A dodge

"We always use the latest models." That sentence means nobody is checking what the latest model broke.

Q.06

What did your last integration actually take?

Ask about the last one, not the typical one. The typical one is a brochure.

A real answer

Weeks, named systems, and the number of people it took on both sides, including yours.

A dodge

"It connects in days." Nothing that touches your identity provider, your CRM, and your data warehouse connects in days.

Q.07

What does it not do?

The single fastest read on vendor quality, and it takes ten seconds to ask.

A real answer

A stated boundary, given without being pushed. A vendor who names the edge has met the edge in production.

A dodge

Scope that widens with every question you ask. A product that does everything has been finished for nobody.

Q.08

Who on our side has to change how they work?

Pilots die at the workflow, not in the model. Somebody's Tuesday has to change or nothing does.

A real answer

Named roles, the new steps in their day, and the training time it takes, stated as a cost of the project.

A dodge

"It fits into your existing workflow." Software that changes nothing about how anyone works also changes nothing about your results.

Q.09

How do we measure this in 90 days?

If this is not agreed before the pilot, the pilot cannot succeed and cannot fail. It will just end.

A real answer

One business metric, agreed now, with the baseline measured before rollout so the comparison exists.

A dodge

Engagement and usage numbers. Usage is not a result. MIT's 95 percent were being used.

Q.10

What happens when we leave?

Every vendor relationship ends. The good ones are designed to.

A real answer

Export formats, contract exit terms, and an honest account of what degrades on the way out.

A dodge

Silence, or a data-retention clause you meet for the first time at signature.

Where these deals go wrong

IN ORDER
F.01

The demo ran on curated data

Yours is not curated. The gap between the demo corpus and your real documents is the single most common place an AI purchase dies, and it is invisible until integration.

F.02

The pilot measured nothing

No baseline, no agreed metric, no end condition. It could not succeed and could not fail, so it ran until the champion changed jobs.

F.03

Nobody's job changed

The software sat next to the workflow instead of inside it. Adoption charts looked fine. The process it was bought to replace kept running in parallel, at full cost.

F.04

It worked at rollout and degraded quietly

Klarna's assistant automated two thirds of its support chats in 2024. By 2025 the company was rehiring humans after quality dropped. The launch press release never gets a sequel. Budget for the sequel.

F.05

It faced customers with no error budget

McDonald's ended its two-year drive-thru voice AI test in 2024 after error clips went viral. Internal tools fail in private. Customer-facing tools fail on TikTok, and in tribunals.

F.06

The integration ate the return

The product worked. The eighteen months of connective work around it cost more than the problem did. This is why Q.02 and Q.06 get asked in the first meeting.

The limit of a page

THE HONEST PART

Here is what this page cannot do for you. You cannot cross-examine a PDF. Vendors rehearse for calls, and a question list loses some of its power the moment the other side has read the same list. The follow-up question, the pause after a vague answer, the reference customer you can ask directly: none of that happens on paper.

The place these questions work best is a room where the vendor's answer can be checked against someone who has actually deployed the thing, in the same hour. That gap between the page and the room is not an accident of format. It is the reason AI 20/20 exists: sessions where the person teaching a use case is the operator running it in production, and the people evaluating vendors get to ask the deployment questions of someone with no commission on the answer. For the record of what has actually shipped, with sources, read the enterprise AI record.

THE NEXT STEP

Ask the questions where they get answered

A vendor's answer is worth what you can check it against. So I write from the systems I actually run: what they do, what they cost me, and where they broke. This guide gets sharper every time one of them surprises me.

 

Keep reading

THE CLUSTER