How to run an AI support pilot without lying to yourself
Most AI pilots are designed to succeed. Here's how to design one that tells you the truth instead — control groups, honest metrics, and the traps that flatter every result.

Most AI support pilots are designed, consciously or not, to succeed. The vendor helps you pick the use case. You choose the friendliest ticket types. You measure before-and-after on a metric that was going to move anyway. Ninety days later you have a deck full of green arrows and no real idea whether the thing works. Then you roll it out, and reality — the full messy queue, the hard tickets, the edge cases — arrives all at once.
The industry's own numbers should make you suspicious of how easy "success" is to declare.
Seventy percent report value in sixty days, where "value" is whatever the respondent decided it meant. That's not a result; it's a mood. Here's how to run a pilot that tells you the truth instead — even when the truth is inconvenient.
The traps, named
The cherry-picked cohort. You pilot on password resets and order status — the tickets AI handles best — then generalise to the whole queue. The pilot succeeds and the rollout disappoints, because you tested the easy 20% and deployed to the hard 100%.
The vendor-run pilot. The vendor configures it, tunes it, and reports the results. Every one of those steps has a thumb on the scale. A vendor running your pilot isn't evaluating their product; they're demonstrating it.
The missing control. Your CSAT went up during the pilot. Compared to what? Without a control group you can't separate the AI's effect from a seasonal dip in volume, a new KB article, or the simple fact that the pilot team knew they were being watched.
The Hawthorne effect. People behave differently when they know they're in the experiment. Your pilot agents are more careful, more engaged, and more forgiving of the tool than a general rollout will ever be.
A pilot that can only succeed isn't a pilot. It's a purchase you haven't admitted to yet.
The design that resists self-deception
1. Write the failure condition first — in advance, in ink. Before the pilot starts, define what result would make you not buy. If you can't name one, you're not running a pilot; you've already decided. "We proceed unless resolution quality drops below X or repeat contacts rise above Y" is a pilot. "Let's see how it goes" is a purchase.
2. Use a real control group. Split comparable volume: some through the AI, some through the current process, at the same time. Random assignment if you can, matched cohorts if you can't. Without a concurrent control, every number you collect is a story, not a result.
3. Pilot the hard tickets too. Include a representative slice of your genuinely difficult work, not just the demo-friendly cases. You need to know how the tool fails on the tickets that matter — because those are the ones customers remember.
4. Measure outcomes, not activity. "Drafts generated" and "tickets touched" are activity. Measure resolution that held (no reopen, no repeat contact within a week), quality-audited by humans, segmented by path. And beware the modelled promises —
— that 30–45% is a potential from an economic model, not a pilot outcome, and it may be the most over-cited "as if realised" number in the field. Your pilot exists to find your real figure, which will be smaller and far more useful.
5. Run it long enough to see the second-order effects. Sixty days catches the productivity bump. It often misses the repeat-contact wave, the trust erosion, and the agent over-reliance that shows up once the novelty wears off. Run long enough for the honeymoon to end.
6. Separate measured from projected. Keep your pilot's measured results apart from the vendor's forecasts.
A forecast about 2029 is a fine reason to run a pilot and a terrible substitute for one. What you want from the exercise is a number you measured yourself, under conditions close to your real operation.
The standard to hold yourself to
The best evidence in this entire field isn't a vendor pilot. It's a controlled study.
You won't match a research team's rigour, and you don't need to. But that's the standard: a real comparison, measured outcomes, honest about where the effect concentrates. Aim at it. A pilot built to survive contact with that standard will sometimes tell you no — and the ability to hear no is the entire point. A pilot that can only say yes has told you nothing except how badly you wanted it to.
When it does say yes — genuinely, measurably — you'll know exactly what you're buying and where it breaks, which is the strongest position a support leader can be in. Take the vendor scorecard into the negotiation, and place whatever you buy on the maturity model so you know what the next stage will require.