QA scorecard & rubric
A weighted, 8-criterion rubric for scoring individual support conversations on a 0–100 scale, with an auto-fail gate — use it to make QA consistent and outcome-focused rather than a checkbox exercise.
A weighted rubric for scoring individual support conversations (ticket, chat, or call). Use it to make QA consistent, defensible, and focused on whether the customer actually got helped — not on whether the rep ticked every box.
How scoring works
- Score each criterion from 1 to 5 (whole numbers).
- Multiply each score by its weight, sum, and divide by 5 to get a 0–100 quality score. (Max raw = sum of weights × 5 = 500;
score = raw / 5.) - Any item on the auto-fail list caps the whole conversation at a failing score regardless of the rubric total — flag it, don't average it away.
- Weights sum to 100. Adjust them for your team, but keep the outcome-heavy tilt (see the note at the end).
The rubric
| # | Criterion | Group | Weight | What it measures |
|---|---|---|---|---|
| 1 | Resolution | Resolution & accuracy | 25 | The customer's actual problem was solved, or moved to a correct next step with a clear owner and timeline. |
| 2 | Accuracy | Resolution & accuracy | 15 | Every fact, instruction, and commitment given was correct and complete. No guessing presented as certainty. |
| 3 | Diagnosis | Resolution & accuracy | 10 | The rep understood the real issue before acting — read the context, asked the right question, didn't solve the wrong problem. |
| 4 | Clarity | Communication & tone | 12 | Response was easy to follow, correctly ordered, and pitched at the right level for this customer. |
| 5 | Tone & empathy | Communication & tone | 10 | Tone matched the customer's state; frustration acknowledged; no canned coldness or false cheer. |
| 6 | Ownership & proactivity | Communication & tone | 8 | Rep took charge, set expectations, and answered the next question before it was asked. |
| 7 | Documentation | Process & compliance | 10 | Notes, tags, and disposition are accurate enough that the next person needs zero re-work. |
| 8 | Policy & security | Process & compliance | 10 | Required steps followed: identity verification, correct escalation path, disclosures, refund/credit limits. |
| Total | 100 |
Group totals: Resolution & accuracy = 50 · Communication & tone = 30 · Process & compliance = 20.
Scoring guide (anchors)
Only the criteria that reviewers disagree on most are anchored here. For the rest, treat 3 = meets the bar, 5 = a model example you'd share in calibration, 1 = would hurt the customer or the business.
Resolution (weight 25)
- 5 — Problem fully solved in-thread, or handed off with a named owner, a committed timeline, and the customer told exactly what happens next. Nothing left for the customer to chase.
- 3 — Problem solved, but the customer had to re-explain, wait without an ETA, or take an extra step that the rep could have handled.
- 1 — Ticket closed while the underlying problem is still live, "solved" with a workaround the rep didn't verify, or bounced back to the customer with no path forward.
Accuracy (weight 15)
- 5 — Everything stated is correct and complete; limits and edge cases called out; any uncertainty is labeled and a real answer is chased down.
- 3 — Core answer is right, but a detail is vague, slightly off, or omitted in a way the customer probably won't hit.
- 1 — Materially wrong information, a made-up feature/policy, or a commitment the business can't keep.
Tone & empathy (weight 10)
- 5 — Reads like a competent human on the customer's side. Frustration is named and defused; warmth is genuine, not scripted.
- 3 — Polite and professional, but generic — could have been sent to anyone about anything.
- 1 — Dismissive, defensive, robotic, or falsely upbeat in the face of a real problem ("Happy to help!" on an outage).
Documentation (weight 10)
- 5 — A colleague could pick this up cold: root cause, actions taken, and current state are all captured; tags and disposition are correct.
- 3 — Notes exist but are thin or use shorthand only the author understands; one field is wrong or blank.
- 1 — No usable notes, wrong disposition, or a tag that will send reporting and routing the wrong way.
Auto-fail list
Any one of these caps the conversation at a fail (typically ≤50), no matter how strong the rest of the interaction was. Reviewers should flag the specific item, not just deduct points.
- Wrong or fabricated information that could cause the customer to lose data, money, or access.
- Identity/security step skipped — released account info or made a sensitive change without required verification.
- Compliance or legal breach — missing required disclosure, mishandled PII, or a promise that violates policy or regulation.
- Ticket closed unresolved or marked solved to beat a metric while the problem is still live.
- Rude, dismissive, or hostile language toward the customer.
- Unauthorized commitment — refund, credit, discount, or timeline the rep had no authority to promise.
- Escalation ignored — clear signal (churn threat, safety issue, exec/legal mention) not routed where it needed to go.
Reviewer checklist (per conversation)
Run this before you finalize a score.
- I read the whole thread, including internal notes, before scoring.
- I checked the outcome against what the customer actually needed, not just what the ticket asked.
- I scored the work, not the rep — same standard I'd apply to anyone.
- I screened every auto-fail item.
- Every score of 1 or 5 has a one-line reason attached.
- My feedback names one thing to keep and one thing to change — specific, quotable, actionable.
Notes on weighting: outcomes over checkboxes
- Weight the result, not the ritual. Resolution and accuracy carry half the score on purpose. A rep who solved the problem cleanly but forgot one tag should outscore a rep who followed every step and left the customer stuck.
- Process weight protects the customer and the business, not the process. The 20 points on documentation and policy exist because bad notes and skipped verification cause future failures — not to reward form-filling. If a process step doesn't map to real risk or a real handoff, drop it from the rubric.
- Auto-fails are the compliance floor — keep them off the main scale. Instead of piling checkbox penalties into the score, gate hard failures through the auto-fail list. This stops a single missed step from silently sinking an otherwise excellent interaction, and stops "checklist completeness" from masquerading as quality.
- Calibrate the anchors, not just the numbers. Most reviewer disagreement comes from Resolution, Tone, and Accuracy. Run periodic calibration on real tickets and tighten the 1/3/5 wording where reviewers split — the rubric is only as consistent as its anchors.
- Review a sample, coach the pattern. Scores are an input to coaching, not the point of it. Track why conversations lose points across a rep or team, and fix the top recurring cause before chasing the average up a point.
Continue exploring