The AI Support Vendor Scorecard
Twelve questions to ask any AI support vendor before you sign — and the answer to each that should make you walk.

Every vendor demo works. That is not a compliment — the demo is the single path the vendor has spent the most time making flawless. The queue is not a demo. So before you sign anything, run the questions the sales deck hopes you won't. Here are twelve, in four groups, each with the answer that should make you walk.
How success is measured
1. Define "resolution" — in writing. The entire business case rests on this one word, and vendors define it generously. Intercom, for example, counts an "assumed resolution" when a customer simply stops replying.
The answer that should scare you: "Resolution is when the customer doesn't come back." Silence is not success. Often it's a customer who gave up and churned quietly.
2. What's the denominator? A "90% resolution rate" over what — all contacts, or the tidy subset the bot chose to answer? Marketed benchmarks are routinely drawn from a best-case slice.
The answer that should scare you: a headline rate with no denominator, or one measured across "top-performing" accounts or industries.
3. Who audits the quality of an automated resolution? A resolved ticket and a well-resolved ticket are different things, and only reading the transcripts tells you which you're getting. Scare: "The system scores itself."
What it costs at scale
4. Model the per-outcome bill at your real volume. Outcome-based pricing sounds perfectly aligned until you multiply it across a year of tickets. Scare: a per-resolution price attached to a fuzzy definition of the billable event — you're paying for the definition, not the outcome.
5. What exactly counts as a billable outcome? Under outcome pricing the vendor is incentivised to classify generously. Scare: a "resolution" is booked the moment the bot replies, before the customer is satisfied.
6. What's the all-in cost, including the escalations it creates? A bot that deflects cheaply but escalates angry, half-handled tickets can raise your cost per genuinely resolved contact. Measure the whole funnel, not the top of it.
What happens when it's wrong
7. Show me the human offramp. When the AI can't help, how fast and how obviously does a person appear? Regulators have already flagged chatbot "doom loops" — customers trapped with no route to a human.
Scare: "It almost never needs to escalate." Everything needs to escalate eventually; a vendor who won't plan for failure hasn't run a real queue.
8. What's the false-confidence rate? How often is it wrong and certain? A hesitant wrong answer is recoverable; a confident one gets sent to your customer. Scare: they have never measured it.
9. What does a customer actually see when it fails? A dead end, silence, or a graceful handoff with the context preserved? The failure experience is the experience most of your unhappy customers will remember.
Whose data, whose model, whose risk
10. Where does our data go — and does it train a shared model? Your transcripts are your customers' data. Scare: vague reassurance about "improving the service for everyone."
11. Can we leave? Ask about export, portability, and what happens to your tuning if you switch. Scare: the tuning that makes it good lives only inside their platform.
12. Who is liable when it tells a customer something wrong? Get the answer before it happens, not after.
A vendor who won't put the definition of 'resolution' in the contract is telling you exactly how the number was made.
Print this. Take it to the demo. A good vendor answers all twelve without flinching — the honest ones are relieved to finally meet a buyer who asks. For where these tools fit in the first place, start with the maturity model; to go deep on the resolution-rate games behind question 1, read what "AI resolved 70% of tickets" actually means.