Root-cause analysis (5 Whys) template
A fill-in 5 Whys worksheet for support quality: take a recurring problem — a reopen spike, a repeat DSAT theme, a ticket type QA keeps flagging — past the symptom to a root cause you can actually change, then assign a countermeasure, an owner, and a recheck date.
Use this when the same problem keeps coming back — a spike in reopens, a repeat DSAT theme, a ticket type QA keeps flagging, or one incident bad enough to warrant a write-up. It is not for one-off mistakes. Run it on patterns, with the people closest to the work in the room.
The 5 Whys came out of the Toyota Production System; Taiichi Ohno used it to trace a stopped machine past the blown fuse to a missing filter. The Lean Enterprise Institute is blunt about one thing: "the specific number five is not the point... keep asking until the root cause is reached and eliminated." Five is a rule of thumb, not a quota. Stop when you hit something you can actually change; keep going if you're still describing a symptom.
Why bother instead of just closing the ticket? Because the ticket isn't the cost. Harvard Business Review's contact-center research found 22% of repeat calls involve downstream issues related to the problem that prompted the original call — even when that first problem was solved. Fix the reply and the customer often comes back anyway. Fix the cause and they don't.
Before you start
- Name one specific, observed problem — a real ticket ID, a metric with a number, a dated incident. Not "CSAT is down," but "billing tickets reopened 31% last month, up from 12%."
- Confirm it's a pattern, not a one-off (≥
[3]occurrences, or one incident above your severity bar). - Pull
[2–3]example tickets so every "why" is grounded in evidence, not opinion. - Get the right people in the room: whoever handles these tickets, plus anyone who owns the
[product / policy / tool]you might land on.
Problem statement
| Field | Fill in |
|---|---|
| The problem, in one sentence | [what is happening] |
| How we know (metric / ticket IDs) | [reopen rate 31%, #…] |
| First seen / how often | [date; N times per week] |
| Who it hits | [customer segment / channel] |
| Rough cost | [volume × cost-per-contact, or CSAT / hours] |
The 5 Whys
Start from the problem statement. Each answer must be a fact you can point to, not a guess. If two causes are both real, branch — write a second chain rather than forcing one line.
| Step | Ask | Answer (evidence, not opinion) |
|---|---|---|
| Why 1 | Why did [the problem] happen? | [ ] |
| Why 2 | Why did answer 1 happen? | [ ] |
| Why 3 | Why did answer 2 happen? | [ ] |
| Why 4 | Why did answer 3 happen? | [ ] |
| Why 5 | Why did answer 4 happen? | [ ] |
| Root cause | The first cause you can actually change | [ ] |
Test the chain: read it bottom-to-top with "therefore." If "no filter → worn shaft → weak pump → poor lubrication → seized bearing → machine stops" reads cleanly, the logic holds. If a step doesn't follow, you skipped a why.
Worked example
| Step | Answer |
|---|---|
| Problem | Refund tickets reopen at 31% |
| Why 1 | Customers reply saying the refund never arrived |
| Why 2 | Reps quote "3–5 days" but the processor now takes 10 |
| Why 3 | The refund macro still says 3–5 days |
| Why 4 | The macro wasn't updated when the processor changed in [month] |
| Why 5 | No one owns keeping macros in sync with vendor changes |
| Root cause | No process links vendor / policy changes to KB and macro updates |
Notice where it landed: not "the rep gave the wrong ETA." Stopping at the person is the most common way this goes wrong — it produces a coaching note and the same tickets next month.
Classify the root cause
Mark where it actually lives. This decides who fixes it.
- Product / engineering — a defect or missing feature → route to product with volume and cost attached
- Knowledge — wrong or missing macro, KB article, or guidance → owner in the KB
- Process — a broken handoff, no owner, no trigger → fix the workflow
- Policy — the rule itself generates the contact → escalate to the policy owner
- Tooling — the system makes the right action hard → tooling backlog
- Skill — a genuine knowledge or technique gap across reps → training, not a one-off scolding
If you landed on "the agent made a mistake," ask one more why: what made the mistake easy to make?
Countermeasure and verification
| Field | Fill in |
|---|---|
| Countermeasure (specific, testable) | [what changes] |
| Owner | [name] |
| Due | [date] |
| How we'll know it worked | [reopen rate < 15% within 30 days] |
| Recheck date | [date] |
- The fix addresses the root cause, not the symptom.
- It has one named owner and a date — "the team will look into it" is not a countermeasure.
- There's a metric and a recheck date. RCA isn't done when you find the cause; it's done when the number moves.
- If the metric hasn't moved by the recheck, you fixed the wrong thing — reopen the analysis.
Common failure modes
- Stopping at the symptom. "The bot gave a wrong answer" is a why-1, not a root cause.
- Stopping at a person. Blame ends the investigation and fixes nothing systemic.
- Single chain when there are two causes. Branch instead of picking the convenient one.
- Opinions dressed as facts. Every answer needs a ticket, a log line, or a number behind it.
- No countermeasure owner. A root cause with no owner is a diary entry.
The point of the exercise: each RCA should retire a slice of avoidable volume for good. If the same theme comes back next quarter, the last analysis stopped one why too early.
Continue exploring