AI pilot readiness checklist
A pre-flight checklist to run before you switch on an AI support pilot, covering baselines, a failure condition written in advance, a control group, hard tickets, a human offramp, data/privacy sign-off, and outcome-based measurement — use it so the pilot can tell you the truth instead of flattering the tool.
A pre-flight checklist to run before you turn on an AI support pilot — chatbot, agent-assist, autoreply, triage, or summarization. Its whole job is to make the pilot capable of telling you the truth: that you set the bar before you saw the results, that you can compare against something, and that a bad outcome is defined in advance instead of argued about after.
Work top to bottom. If you can't check a box in Phase 1 or 2, you're not ready to launch — you're ready to do the missing work. Most failed pilots don't fail at the model; they fail because nobody wrote down what "working" meant while they still could be honest about it.
Phase 0 — Frame the pilot
Decide what this pilot is actually testing before you scope anything else. A pilot answers one question well or three questions badly.
- One sentence, written down: "We are testing whether [AI capability] can [specific outcome] for [specific ticket type / queue] without [specific harm]."
- Scope is narrow. One channel, one queue or intent, one clearly bounded ticket type — not "support" in general.
- A named owner is accountable for the go/no-go call (not a committee, not the vendor).
- A time box is set: a start date, an end date, and a minimum ticket volume the pilot must reach before anyone reads results (enough to not be noise — for most teams, hundreds of AI-touched tickets, not dozens).
- The vendor is not the scorekeeper. You decide what counts as success and you hold the measurement.
Phase 1 — Know your baseline
You cannot prove improvement against a number you don't have. Pull these for the same queue and ticket type the pilot will touch, over a representative recent window (e.g. the last 4–8 weeks), before launch.
| Metric | Why it's here | Have it? |
|---|---|---|
| True resolution rate — issue actually solved, not just ticket closed | The number the pilot will be tempted to game | ☐ |
| Reopen rate (or repeat-contact within 7 days) | Catches "closed but not solved" | ☐ |
| CSAT / customer-reported outcome for this ticket type | Guards against speed-at-the-cost-of-help | ☐ |
| Full resolution time (first contact → truly resolved) | Not just first-response time | ☐ |
| Escalation / transfer-to-human rate | Your baseline offramp load | ☐ |
| Cost per resolved contact (or handle time as a proxy) | So savings claims are real, not assumed | ☐ |
| Volume & mix for the queue | So you can tell if the pilot got the easy tickets | ☐ |
- Every metric above is defined in writing — including exactly when a ticket counts as "resolved."
- You know the current baseline variance, not just the average (last quarter's week-to-week range), so you can tell a real change from noise.
- Baselines are segmented by the ticket types the pilot will and won't touch, so a change in mix can't masquerade as a change in performance.
Phase 2 — Define failure before you launch
This is the phase teams skip, and it's the one that makes a pilot honest. Write the kill conditions while you still have nothing to defend.
Kill criteria (written in advance)
- A failure condition is written down and signed off before launch — a specific, measurable line that, if crossed, ends or pauses the pilot. Example wording: "We stop if reopen rate on AI-handled tickets exceeds baseline + 20%, or CSAT drops more than 3 points, sustained over any 3-day window."
- At least one guardrail metric protects quality so it can't be traded for speed or deflection (e.g. reopen rate, CSAT, escalation-after-AI rate).
- A hard stop / circuit breaker exists: who can pause the pilot, how fast (target: minutes, not a meeting), and what the rollback looks like.
- "Deflection" and "contained" are explicitly not success on their own. A ticket the customer abandoned in frustration is not a win — decide now how you'll tell a real resolution from a customer giving up.
Control group
- A control group is defined: comparable tickets handled the current way, over the same window, so you're measuring the AI — not a good week, a seasonal dip, or a coincident process change.
- Assignment to AI vs. control is as random as your stack allows (or matched on ticket type, channel, and customer segment). Don't let the AI cherry-pick.
- You've decided how to handle spillover (e.g. AI-drafted replies a human edits) so "AI" and "human" buckets don't blur.
Representative ticket mix
- Hard, ambiguous, and angry tickets are included — not just the FAQ-shaped ones. A pilot fed only easy tickets tells you nothing about production.
- The pilot's ticket mix matches the real queue's mix (edge cases, multi-issue threads, non-native-language, low-context requests), or you've documented exactly what you excluded and why.
- A small set of known-nasty test tickets (past escalations, known failure modes, prompt-injection / jailbreak attempts) is queued to run against the system before and during the pilot.
Phase 3 — Design the human offramp
Every AI support flow needs a way out that a real person controls. Design it now, not after the first angry escalation.
- There is a clear, always-available path to a human — the customer can reach one without fighting the bot or guessing a magic word.
- Escalation triggers are defined: low confidence, repeated failure to resolve, detected frustration/sentiment, explicit human request, and any high-risk topic (billing disputes, cancellations, safety, legal, vulnerable customers).
- Context transfers with the handoff. The human inherits the full conversation and any AI actions taken — the customer never re-explains from zero.
- Staffing covers the offramp. Someone is actually available to catch escalations during pilot hours; the human queue won't quietly overflow.
- Reps know what's happening. The team whose tickets are in the pilot has been briefed on what the AI does, how to take over, and how to report a bad AI response in one click.
- A fast feedback loop exists for reps to flag wrong/harmful AI outputs, and someone owns triaging those daily.
Phase 4 — Clear data & privacy
Do this before a single real customer message reaches the model. This is the box that turns into a headline if you skip it.
- Data flow is mapped and cleared: you know exactly what customer data leaves your systems, to which vendor/sub-processors, in which region, and it's covered by a DPA and your privacy policy.
- Training / retention terms are confirmed in writing — whether the vendor trains on your data, how long prompts and outputs are retained, and how to delete them.
- PII handling is decided: redaction/masking where feasible, and an explicit call on sensitive categories (health, financial, children's data) that may be out of scope entirely.
- Access controls and logging are in place: who can see AI conversations, and every AI action is logged for audit and incident review.
- Security/legal/privacy signed off — named approver, on record — and you can honor a data-subject deletion request that includes AI-processed content.
- Customer-facing disclosure is decided where required (that they're talking to / being assisted by AI), consistent with your obligations and region.
Phase 5 — Measure on outcomes, not activity
Instrument the pilot so the results can't be spun. Set this up before launch, not while you're staring at a dashboard someone else built.
- Success is defined as an outcome, not an activity. The headline metric is problem actually resolved (and staying resolved), not deflection rate, containment, response speed, or "tickets touched by AI."
- The same guardrails from Phase 2 (reopen rate, CSAT, escalation-after-AI) are on the scoreboard next to any efficiency metric — a speed or cost win that moves these the wrong way is not a win.
- Attribution is honest: you can separate tickets the AI truly resolved from tickets it deflected, delayed, or handed to a human who did the real work.
- A sample of AI-handled conversations is QA'd by a human against the same rubric you use for reps — automated metrics alone will lie to you.
- You measure the full journey, including reopens and downstream contacts, over a window long enough to catch "solved on Monday, back on Thursday."
- The pilot's results are compared against the control group, on the same metrics, over the same period — not against last year or against the vendor's benchmark.
Phase 6 — The go / no-go gate
Before launch, agree what each verdict means. Fill this in on day one; you're only allowed to read it at the end.
| Outcome | Decision | Written before launch? |
|---|---|---|
| Beats control on the primary outcome and holds guardrails | Expand — wider queue, more volume, keep measuring | ☐ |
| Mixed — helps on some ticket types, not others | Narrow & re-run — scope to where it worked, retest | ☐ |
| Any kill criterion crossed, or no real gain over control | Stop — document why, keep the baseline, don't relaunch the same thing | ☐ |
- The decision rule is agreed and signed off before launch, so results can't be renegotiated afterward.
- A short, honest write-up is committed to regardless of outcome — what you tested, what happened, what you'd change — so a "no" is a learning, not a loss.
The seven questions this checklist forces you to answer
If you're short on time, these are the load-bearing ones. A pilot that can't answer all seven isn't ready.
- What's the baseline? — you have the numbers this queue does today.
- What counts as failure? — written, measurable, signed off before launch.
- Compared to what? — a real control group, not a good week.
- Does it handle the hard stuff? — nasty and ambiguous tickets are in the pilot.
- How does a customer reach a human? — a real offramp someone staffs.
- Is the data cleared? — mapped, DPA'd, and privacy/legal signed off.
- How do you know it worked? — measured on resolved outcomes, not deflection.
Continue exploring