Escalation matrix
A ready-to-fill escalation matrix that maps issue types and severities to a named owner, target response time, and the exact next tier or specialist — use it when your queue keeps stalling on tickets nobody clearly owns.
An escalation matrix answers one question fast, under pressure: this ticket, right now — who owns it, how long do I have, and where does it go next? Fill in the tables below with your own tiers, tools, and names, then pin it where the queue can see it. If a rep has to ask someone what to do with a ticket, the matrix has a gap. Close the gap.
How to read this
- Tier = a level of ownership, not a person. T1 = frontline/general queue. T2 = product or senior specialists. T3 = engineering / the people who can change the product. "Specialist" = a named function (Billing, Security, Trust & Safety) you route to directly, skipping tiers.
- Target response = time to a human acknowledgement with a real next step, not an autoreply. Measure from ticket creation, not from when someone opened it.
- Owner = the single name accountable until resolved or formally handed off. Ownership never becomes "the queue." A ticket with no name is an unowned ticket.
1. Severity definitions
Severity is about impact and blast radius, not how loudly the customer is asking. A calm email reporting a data leak is Sev1. An angry ticket about a typo is Sev4. Decide severity from these rules, not from tone.
| Sev | Name | Definition (any one criterion qualifies) | Target first response | Update cadence | Auto-notify |
|---|---|---|---|---|---|
| Sev1 | Critical | Full outage, data loss/exposure, security breach, payment system down, or a legal/safety risk. Affects all or many customers, or a single customer with contractual severity. No workaround. | 15 min, 24/7 | Every 30 min until mitigated | On-call lead + eng on-call + Duty Manager |
| Sev2 | High | Major feature broken or severely degraded; affects a segment of users or a key account; workaround is painful or partial. Money or SLA is on the line. | 1 business hour | Every 2 hours | Team Lead + owning specialist |
| Sev3 | Medium | Single-user or non-blocking issue with a reasonable workaround. Annoying, not stopping work. The bulk of the queue. | 1 business day | Daily until resolved | Owning tier only |
| Sev4 | Low | Cosmetic, docs, feature request, "how do I…". No workaround needed because nothing is broken. | 3 business days | On status change | None |
Severity is a dial, not a label. Re-grade the moment reality changes: if a Sev3 turns out to hit 200 accounts, it's a Sev2 now — bump it and re-notify. Anyone can raise severity; only the owning Team Lead (or above) can lower it, and they note why.
2. The escalation matrix
Map each issue type to its default severity, its first owner, and the named next stop. "Escalate to" is a destination, not "ask around."
| Issue type | Default sev | First owner | Target response | Escalate to → | Trigger to escalate |
|---|---|---|---|---|---|
| Login / auth broken (many users) | Sev1 | T1 → Duty Manager | 15 min | Identity/Platform on-call | Confirmed >1 account or SSO down |
| Suspected security / data exposure | Sev1 | Security on-call (direct) | 15 min | Security Lead + Legal | On first credible signal — do not triage in queue |
| Payments / billing outage | Sev1 | Billing specialist (direct) | 15 min | Payments eng on-call | Charges failing or double-charging |
| Data loss / corruption | Sev1 | T2 | 15 min | Eng on-call (T3) | Any irreversible data change |
| Key/named account fully blocked | Sev2 | Account's named CSM/T2 | 1 hr | Escalation Manager | Account can't do core job |
| Feature broken, workaround exists | Sev2 | T2 | 1 hr | Product on-call | Reproduced + >5 tickets |
| Bug, single user, workaround | Sev3 | T1 | 1 business day | T2 (with repro steps) | T1 can't reproduce or fix in 2 touches |
| Billing question / dispute | Sev3 | Billing specialist | 1 business day | Finance | Refund > your approval limit |
| Legal / compliance / GDPR request | Sev3 | T1 → Legal (direct) | 1 business day | Legal + DPO | Any DSAR, subpoena, or "my lawyer" |
| Abuse / threat / self-harm mention | Sev2 | Trust & Safety (direct) | 1 hr | T&S Lead + Duty Manager | On first mention — never sit on it |
| How-to / config help | Sev4 | T1 | 3 business days | T2 or docs | Only if genuinely undocumented |
| Feature request | Sev4 | T1 | 3 business days | Product board | Log it; never escalate live |
Copy the row shape, delete our examples, and write yours. Every row needs a named human or on-call rota in the owner and escalate-to columns — a team name with no rotation attached is where tickets go to die.
3. What must never sit in the general queue
Some tickets are unsafe to leave in first-in-first-out. These get pulled the instant they're identified and routed straight to a named owner, regardless of position:
- Security / data exposure — breach, leaked credentials, "I can see another customer's data." Straight to Security on-call.
- Legal-bearing language — "lawyer," "sue," "GDPR/CCPA," "data subject request," regulator, subpoena, chargeback dispute. Straight to Legal.
- Trust & Safety — threats, harassment, illegal content, or any mention of self-harm. Straight to T&S, with a defined duty-of-care script.
- Payment failures at scale — failed, double, or missing charges. Straight to Billing + Payments eng.
- Anything already Sev1 or Sev2 — by definition it has a named owner and a clock; it does not wait in line.
- VIP / contractual-SLA accounts — routed by account tag to their named owner before general triage.
- Anything aged past its SLA — an about-to-breach ticket is an escalation, not a queue item.
Rule of thumb: if getting it wrong causes irreversible harm — to a person, to data, or to the company's legal position — it does not belong in the general queue. When in doubt, pull it and escalate; over-escalating one ticket is cheaper than sitting on the wrong one.
4. When to escalate (the trigger checklist)
Escalate the moment any of these is true. Don't wait for the customer to ask twice.
- I've made two real attempts and the issue isn't resolving (the "two-touch rule").
- The fix requires access or authority I don't have (a config, a refund over my limit, a code change).
- I can't reproduce it and the customer can.
- The blast radius is growing — more tickets on the same root cause are arriving.
- The ticket is approaching its SLA with no clear path to resolve.
- It matches a "never in the queue" category above.
- The customer is a named/contractual account and is blocked.
- My gut says this is bigger than it looks. Trust it — flag it and let a lead decide.
Escalating is not failure. Sitting on something you can't fix is.
5. The handoff — what every escalation must include
A handoff without context is just moving the problem. Before you pass a ticket up or across, the escalation note carries:
- Ticket link and customer/account (with SLA tier).
- Severity and why (which criterion from §1).
- What's happening in one plain sentence — impact, not symptoms.
- Steps to reproduce, or "can't reproduce" + what you tried.
- What you've already done — the two touches, so no one repeats them.
- What you need from the next owner — the specific ask.
- Customer's last message + expectation you set (e.g. "told them we'd update by 3pm").
- @named owner tagged and acknowledgement confirmed — escalation isn't done until someone says "I've got it."
Ownership transfers on acknowledgement, not on send. If no one has said "got it," you still own it. No silent handoffs.
6. Roles referenced (define these for your org)
| Role / rota | Who it is here | How to reach | Hours |
|---|---|---|---|
| Duty Manager | e.g. rotating lead | #duty-manager / phone | 24/7 |
| Eng on-call (T3) | PagerDuty rota | page | 24/7 |
| Security on-call | _____ | security@ / page | 24/7 |
| Billing / Payments | _____ | #billing | Business hrs |
| Legal / DPO | _____ | legal@ | Business hrs |
| Trust & Safety | _____ | #trust-safety | 24/7 |
| Escalation Manager | _____ | _____ | Business hrs |
Keep it honest
- Review the matrix monthly: any ticket that got escalated to the wrong place, or bounced back, is a row that needs rewording.
- Track how often each trigger fires. A trigger that never fires is probably being ignored; one that fires constantly is a product problem to fix upstream, not a queue to staff.
- If two people would route the same ticket differently, the matrix is ambiguous. Fix the wording, not the people.
Continue exploring