Most "how to choose an AI vendor" checklists ask about security, pricing and
years of experience. Every vendor you talk to will answer all three
favorably, because no vendor is going to tell you their security posture is
weak or their pricing is unreasonable. The checklist filters out nobody. It
exists to make the buyer feel diligent, not to actually surface risk.
The questions that expose real risk are the ones a vendor can't comfortably
prepare a polished answer for in a sales call, because answering them well
requires having actually shipped something and lived with the consequences.
Here are four, and the specific failure mode each one is designed to catch.
"Walk me through a project where you said no to a client's requested scope"
Ask this and watch how long it takes them to come up with a real example.
A vendor who has shipped enough production AI projects has scars — a
feature they talked a client out of building because it wasn't going to
work reliably, a scope cut they insisted on because the client's timeline
didn't match the actual complexity. If the answer is vague, generic, or
takes visible effort to construct, that's a signal: either they haven't
shipped enough to have accumulated this experience, or their business model
rewards agreeing to whatever the client asks for rather than pushing back
when it's the wrong call.
A vendor that always says yes is a vendor whose v1 will be bloated with
features that don't work well, because nobody on their side is incentivized
to tell you a good idea is a bad first version.
"What would you leave out of a v1?"
This tests the same muscle from the other direction. A team with real
delivery discipline has a clear, immediate answer — they know exactly which
parts of an AI feature are hard to get right the first time and which parts
are safe to defer. A team without that discipline will either try to build
everything at once, or will answer with something evasive like "it depends
on your requirements" without following it with the actual factors it
depends on.
The specific risk this exposes: teams that don't know what to cut from a v1
usually ship late, over budget, or with an unreliable full-scope version
instead of a reliable narrow one.
"How do you measure whether the AI is actually working after launch, not just at demo time?"
A demo is curated. Production is not. Ask specifically how they measure
real-world performance once the system is live — an evaluation set run on
production traffic, a defined accuracy or resolution metric tracked over
time, a process for catching when a model update or a shift in user
behaviour degrades performance. If the answer is "we monitor it closely" or
"we're always iterating" with no specifics about what's actually measured,
that's a team that will find out about a regression from an angry customer
instead of a dashboard.
This is the single clearest tell for whether a vendor has evaluation
discipline, because building an evaluation harness is unglamorous work that
teams only do if they've been burned by not having one.
"Who owns a bug found in month four, after launch?"
AI systems degrade in ways traditional software doesn't — a model provider
updates their model, real user input drifts from what the system was tuned
on, an edge case that didn't show up in testing appears at scale. Ask
directly who is responsible for fixing something that breaks after the
initial engagement ends, and get the answer in writing before you sign, not
after.
Vague answers here — "we'd discuss it at the time," "that would be a new
engagement" without a clear definition of what counts as new — predict a
specific and expensive outcome: a production bug that sits unfixed for
weeks while you and the vendor argue about whose scope it falls under.
What a good answer actually sounds like
Contrast the vague version with what a team that's actually lived through
this sounds like when they answer. On scope, a good answer names the
feature, the reason it wouldn't have worked, and what shipped instead — "the
client wanted the assistant to auto-approve refunds over a threshold; we
scoped it to draft the approval and route it to a human for anything over
₹5,000, because the failure cost of an autonomous wrong approval was higher
than the time saved." On post-launch ownership, a good answer names a
specific arrangement — a support retainer, a defined bug-fix SLA, a clear
line between "bug in what we built" and "new feature request" — not a
promise to "figure it out together." Specificity is the signal. A team that
has actually done this has the specifics on hand; a team that hasn't will
speak in the same register the whole call, regardless of which of the four
questions you ask.
Why this works better than a feature or credential checklist
Every question above is rooted in a real failure mode we've seen play out on
projects, not a hypothetical concern. A vendor that can't answer them well
isn't necessarily incompetent — they might genuinely be new to shipping AI
products at production scale, which is a legitimate stage to be at. But you
should know that going in, and price the engagement, and your own oversight
of it, accordingly.
None of this replaces checking security and pricing — do that too. It just
won't tell you the thing that actually determines whether the project
succeeds after signature: whether this team has shipped past the demo
stage, knows how to tell if their own work is actually working, and sticks
around when it isn't. Ask the four questions above in your next vendor
call and see how the answers compare.