AI support vendor RFP template
A request-for-proposal template you send to AI support vendors that forces their resolution, data, pricing, and accountability claims into writing before any demo — so you compare answers on paper instead of buying the best pitch.
This is the document you send to AI support vendors — not the scorecard you use to rate them internally. An RFP (request for proposal) does one job: it forces every vendor to put the same claims in writing, in their own words, against the same questions, before anyone runs a demo or quotes a discount. A demo shows you the happy path. Written answers to pointed questions show you where the product actually is.
AI support vendors sell outcomes — "resolves 70% of tickets," "reduces cost per contact," "deflects at scale." Those numbers dissolve the moment you ask how it's counted. The whole point of this RFP is to pin the definitions down on paper, where they can't be walked back in a meeting where everyone is staring at a slide.
Send it to every shortlisted vendor with the same deadline, the same context, and the same word limits. Then score the written responses side by side using your vendor scorecard — only invite the survivors to demo. Replace everything in [brackets] before you send.
How to run this RFP
- Fill the context block below once — every vendor gets the identical version
- Set a response deadline (
[2–3 weeks]) and a format: written answers to the numbered questions, in order, with the word limits kept - Send to
[3–5]shortlisted vendors; do not send to a long list — you will not read 12 responses properly - Name a single point of contact for vendor questions so answers are consistent across bidders
- Score responses before any live demo; the demo confirms the paper, it doesn't replace it
- Tell vendors up front: "Coming soon," "on the roadmap," and "we can build that" are read as no. Answer for what ships today.
Context we're giving you (fill once)
Give every vendor enough to quote real numbers instead of a starter tier. Vague context gets vague, un-comparable proposals back.
| Field | Our details |
|---|---|
| Company / industry | […] |
| Monthly ticket / contact volume | […] |
| Channels in scope | [email / chat / voice / social / in-app] |
| Current help desk & stack | [Zendesk / Salesforce / …, CRM, telephony] |
| Languages we support | […] |
| Ticket types this AI would touch | [e.g. billing, how-to, order status] |
| Ticket types explicitly out of scope | [e.g. cancellations, disputes, safety] |
| Knowledge sources | [public KB, internal KB, past tickets, product docs] |
| Deployment mode we want | [autonomous / agent-assist / triage / summarization] |
| Timeline & must-be-live date | […] |
Section 1 — Resolution & performance (the load-bearing section)
This is where AI vendors' headline numbers live or die. Make them define the metric before they quote it.
1.1 State your exact definition of a "resolution" (or "deflection," "containment," "autonomous resolution" — whichever you report). When precisely is a ticket counted as resolved? Does a customer who stops replying without confirming count? Does an abandoned chat count?
1.2 For a customer of our size, ticket mix, and knowledge quality, what resolution rate do you expect — and what did comparable customers actually see in month 1 vs. month 6? Provide the methodology behind any benchmark figure.
1.3 How do you distinguish a genuinely resolved ticket from one the customer abandoned in frustration? What do you measure to tell them apart?
1.4 What is your measured accuracy / hallucination rate, how do you measure it, and what happens when the model is wrong or unsure — does it guess, hedge, escalate, or say "I don't know"?
1.5 What are your guardrails against fabricated policies, prices, or promises? Who is accountable when the AI states something incorrect to our customer?
Why 1.5 has teeth: when Air Canada's chatbot gave a customer wrong information about bereavement fares, a tribunal held the airline — not the chatbot "as a separate legal entity" — liable for the loss (CBS News). The vendor's wrong answer becomes your problem, in front of your customer. Get the accountability terms in writing.
A complete answer includes:
- A precise, testable definition of "resolved" — including how non-responders and abandonments are treated
- Real month-1 and month-6 numbers from a comparable customer, with methodology
- A stated accuracy figure and how it's measured
- A defined behavior for low confidence (not "it's very accurate")
Section 2 — Data, privacy & training
Answer before a single real customer message reaches your model.
2.1 Exactly what customer data leaves our systems, to which sub-processors, in which region? Provide your current sub-processor list.
2.2 Do you train models on our data? If so, on what, for how long, and how do we opt out? State it in writing, not "we can discuss it."
2.3 What are your prompt and output retention periods, and how do we delete data — including AI-processed content — on request?
2.4 How do you handle PII and sensitive categories (health, financial, children's data)? Redaction, masking, or exclusion?
2.5 Provide your DPA and confirm compliance with [GDPR / CCPA / our region].
A complete answer includes:
- A named sub-processor list and data-flow description
- An explicit, written training-opt-out and retention/deletion terms
- A ready-to-sign DPA
Section 3 — Security & compliance
3.1 Provide your current SOC 2 Type II (or ISO 27001) report. Which underlying model provider(s) do you use, and what are their data terms?
3.2 How do you defend against prompt injection and jailbreaks — e.g. a customer message that tries to make the agent issue a refund or leak another customer's data? What testing have you done?
3.3 Is every AI action and conversation logged for audit? Can we export those logs?
3.4 Do you offer SSO (SAML/OIDC) and role-based access — included, not upsold?
A complete answer includes:
- A SOC 2 report on request and named model providers with their terms
- A concrete injection/jailbreak defense, not "our model is safe"
- Full, exportable audit logging
Section 4 — Knowledge & integration
The AI is only as good as what it reads and what it can do.
4.1 What knowledge sources does the agent answer from, and how does it stay current when our KB changes? How do you keep it from answering out of stale or wrong articles?
4.2 What actions can the agent take in our systems (look up an order, issue a refund, reset a password), and how are those scoped and permissioned?
4.3 List your native, supported integrations with [our help desk, CRM, telephony] — not "we have an API." Which direction does data flow, and in near-real-time?
4.4 How do we test the agent against our own tickets in a sandbox before it touches a live customer?
A complete answer includes:
- A clear knowledge-refresh mechanism and stale-content handling
- Scoped, permissioned actions — not unrestricted system access
- Named, supported connectors for our actual stack
- A real sandbox we can load our own tickets into
Section 5 — Human oversight & control
5.1 How does a customer reach a human, and what triggers an escalation (low confidence, repeated failure, detected frustration, explicit request, high-risk topic)?
5.2 Does the full conversation and any AI actions transfer to the human, so the customer never re-explains from zero?
5.3 Is there a kill switch — who can pause the AI, how fast, and what's the rollback? Can we set confidence thresholds and topics the AI must never handle?
5.4 Can our team QA AI conversations against the same rubric we use for reps, and correct/override the agent?
A complete answer includes:
- A defined, always-available human offramp with context transfer
- A fast circuit breaker and configurable no-go topics
- Human QA and override on AI output
Section 6 — Commercial & pricing
Make the billing model unambiguous. This is where usage-based AI pricing quietly compounds.
6.1 State your pricing model in full: per resolution, per seat, per usage, or hybrid — and exactly which events are billable. Is a customer who stops replying billed? Is an escalated-to-human ticket billed?
6.2 Quote total year-one cost for our real volume, all fees included (implementation, integrations, overage). Model it at 2× our volume and at renewal.
6.3 Are there volume caps, overage rates, and minimums? What's the uplift cap at renewal, in writing?
6.4 What's the exit ramp — data export, offboarding help, no auto-renew trap?
Why 6.1 has teeth: Intercom's Fin bills $0.99 per outcome, and an assumed resolution — a customer who simply disengages for 24 hours after the agent's last answer — is billable (Fin pricing). If you don't define the billable event, the vendor's definition is the one you pay for. A ticket the customer gave up on can still cost you a dollar.
A complete answer includes:
- A precise billable-event definition — including non-responders and human escalations
- All-in year-one cost at our volume, modeled at 2× and at renewal
- Written caps, overage rates, uplift limits, and a clean exit
Section 7 — Support, SLAs & roadmap
7.1 What's your uptime SLA and incident-response process when the AI misbehaves in production?
7.2 How much notice do we get before you change or deprecate the underlying model, which can shift the agent's behavior overnight?
7.3 What's your implementation plan, owner, and timeline for a customer like us?
7.4 What shipped in the last two quarters, and what slipped?
A complete answer includes:
- A written uptime SLA and incident process
- A model-change notice commitment
- A named implementation owner and dated plan
Section 8 — Proof
- Provide 2 reference customers of our size and mix we can speak to without you on the call
- State your pilot terms: length, exit, and success metrics we define (see running an AI pilot honestly)
- Confirm you'll run the pilot on our hard, ambiguous, and angry tickets — not a curated demo set
Mandatory requirements (pass / fail)
A "no" on any of these ends the evaluation, regardless of how strong the rest of the proposal reads.
| # | Requirement | Meets? |
|---|---|---|
| M1 | Precise, written definition of "resolution" and every billable event | ☐ |
| M2 | Written training-opt-out, retention, and deletion terms + signable DPA | ☐ |
| M3 | Current SOC 2 Type II (or ISO 27001) on request | ☐ |
| M4 | Always-available human offramp with full context transfer | ☐ |
| M5 | Kill switch + configurable no-go topics | ☐ |
| M6 | All-in year-one pricing at our real volume, with renewal uplift cap | ☐ |
| M7 | Sandbox to test against our own tickets before go-live | ☐ |
| M8 | Two comparable references reachable without the vendor present | ☐ |
Response rules for vendors (put these in the send-out)
- Answer every numbered question, in order, with the section numbers
- Keep each answer under
[150]words; attach evidence (reports, sample logs) separately - Answer for what ships today — mark anything not yet generally available as "roadmap," and we will score it as absent
- Provide pricing for our stated volume, not a starter tier
- One named contact for our follow-up questions
Evaluation & timeline
| Stage | Date |
|---|---|
| RFP sent to shortlist | [date] |
| Vendor questions due to us | [date] |
| Written responses due | [date] |
| Paper scoring complete | [date] |
| Demos (survivors only) | [date] |
| Pilot with finalist | [date] |
| Decision | [date] |
Red flags in the responses
Any one is a reason to dig; stacked up, they predict a bad year.
- Can't define "resolution" without hedging, or defines it as "customer stopped replying"
- Dodges the training-opt-out or retention question ("we can discuss during contracting")
- Answers security with adjectives, not documents (no SOC 2 on request)
- No named accountability for a wrong AI answer to a customer
- Pricing is usage-based with no cap, or the billable event is left fuzzy
- Key answers are "coming soon" or "we can build that for you"
- Won't let you test on your own tickets before go-live ("we'll cover that in the POC")
- References are hand-picked and only available with the vendor on the call
A confident demo is not a written answer. If a vendor won't commit a definition, a price, or a data term to paper in an RFP, they will not commit to it at renewal either. The responses to these questions — not the pitch — are what you're buying.
Continue exploring