QA calibration session agenda
A 60-minute agenda for a recurring QA calibration session where reviewers blind-score the same tickets, surface disagreement, and resolve to a shared standard; use it to keep your QA scores consistent across reviewers.
Run this every two weeks with the same core group of reviewers. The goal is not to grade agents — it's to prove that your reviewers would give the same ticket the same score. Where they wouldn't, you have a rubric problem, not a people problem. Fix it in the room.
Before you start
Everyone needs the same rubric version open, the same tickets in front of them, and a way to score without seeing each other's numbers.
Facilitator prep (day before, ~20 min)
- Pick 3–4 recently closed tickets that are instructive, not obvious. Aim for a mix: one clean resolution, one where policy is ambiguous, one emotionally charged, one you personally scored and weren't sure about.
- Strip agent and reviewer names from each ticket (calibrate the work, not the person).
- Paste the current rubric version number at the top of the scoring sheet (e.g.
Rubric v4 — updated 2026-06-11). - Build a scoring sheet with one hidden column per reviewer so nobody anchors on anyone else's score.
- Name a scribe (not the facilitator) to own the disagreement table live.
- Send the tickets + rubric out, but ask people not to score before the session — blind means blind and fresh.
In the room
- 3–6 reviewers. Fewer than 3 isn't calibration; more than 6 and nobody talks.
- Cameras on, scoring sheet shared, chat open for pasting exact wording.
- Ground rule: you defend the score, not the agent. "I gave it a 1 because…" not "well, Sam was probably busy."
The rubric you're calibrating on
This is the example scale below — swap in your own dimensions, but keep the anchors this concrete. Every dimension is scored 0 / 1 / 2, plus two independent pass/fail flags that can zero the whole ticket.
| Dimension | 0 | 1 | 2 |
|---|---|---|---|
| Resolution — did the reply solve the actual problem? | Wrong problem, or punted | Solved, but customer must do avoidable extra work | Fully solved in this reply |
| Accuracy — facts, steps, and policy correct? | A materially wrong statement | Minor slip, not misleading | Correct throughout |
| Completeness — every question answered, expectations set? | Left a direct question unanswered | Answered the ask, missed a next-step or timeline | Answered everything + set expectations |
| Tone & empathy — matched the customer's state? | Robotic, defensive, or mismatched | Fine but generic | Genuinely human, matched the moment |
| Clarity — easy to read and act on? | Confusing or jargon-heavy | Readable, slightly cluttered | Clean, scannable, plain |
Pass/fail flags (either one = ticket fails, regardless of score):
- Compliance — required disclosure given, correct disposition/notes, no policy breach.
- No harm — no rude, misleading, or trust-damaging statement to the customer.
The 60-minute agenda
| Time | Minutes | Segment |
|---|---|---|
| 0:00–0:03 | 3 | Kickoff: purpose, and the one rule (defend the score, not the person) |
| 0:03–0:08 | 5 | Rubric refresher — read the 0/1/2 anchors aloud, note anything already fuzzy |
| 0:08–0:28 | 20 | Blind scoring — silent, independent, ~5 min/ticket, into hidden columns |
| 0:28–0:33 | 5 | Reveal all scores at once; scribe flags every gap that hits the threshold |
| 0:33–0:50 | 17 | Work the disagreements — dimension by dimension, hardest ticket first |
| 0:50–0:57 | 7 | Resolve each to an agreed score; capture rubric edits as you go |
| 0:57–1:00 | 3 | Assign owners for rubric changes + set next session |
0:08–0:28 — Blind scoring
- Score all 3–4 tickets silently and independently. No chat, no faces, no "wait, is this a 1 or a 2?" out loud.
- Score every dimension even when unsure — a guess reveals more disagreement than a blank.
- Set both pass/fail flags per ticket.
- Facilitator watches the clock: ~5 minutes a ticket, then move on.
0:28–0:33 — Reveal and flag
Reveal all columns at once. The scribe flags any line that meets the disagreement threshold:
Surface it if the score spread on a dimension is ≥ 2 points, or a pass/fail flag disagrees between reviewers.
A 1-vs-2 split is noise; leave it. A 0-vs-2 split, or one reviewer failing a ticket others passed, is exactly what this meeting exists to resolve.
0:33–0:50 — Work the disagreements
Take the ticket with the widest spread first. For each flagged dimension:
- The highest and lowest scorer each say, in one sentence, why — quoting the ticket.
- Name the root cause (see the table). This is the whole game: most disagreement is the rubric being vague, not people being wrong.
- Land on one agreed score. If you genuinely can't, that's a rubric gap — log it and move on; don't burn 10 minutes forcing consensus.
Disagreement tracking table
The scribe fills this live. It is the output of the session.
| Ticket | Dimension | Spread (low→high) | Root cause | Agreed score | Rubric action |
|---|---|---|---|---|---|
| #4471 | Completeness | 0 → 2 | Rubric ambiguity — "set expectations" undefined | 1 | Add anchor: must state a timeline or next step |
| #4471 | Compliance flag | Fail vs Pass | Missing context — reviewer didn't see the refund note | Pass | None (context, not rubric) |
| #3902 | Tone | 0 → 2 | Genuine judgment call — firm vs cold | 1 | Add example of acceptable "firm" wording |
| #3902 | Accuracy | 1 → 1 | No disagreement — omit | — | — |
Root-cause tags (pick one per row):
- Rubric ambiguity — the wording lets two honest people score differently. → Rubric action.
- Missing context — someone scored without info the others had. → No rubric change; fix how tickets are prepped.
- Misread — a reviewer simply missed a line. → No change; note it.
- Genuine judgment call — the rubric is clear, values differ. → Add a worked example to anchor it.
0:50–0:57 — Resolve and update the rubric
- Every flagged line has an agreed score in the table.
- Every Rubric ambiguity and Judgment call row has a concrete edit — new anchor wording or a worked example, not "clarify this."
- Bump the rubric version and date the change.
- The scored tickets become reference examples — file the agreed scores where reviewers can find them.
0:57–1:00 — Close
- Assign an owner and a date to each rubric edit (default: facilitator, within 3 days).
- Log two health numbers so you can watch them trend:
- Agreement rate — dimensions scored within 1 point ÷ total dimensions scored.
- Open rubric gaps — rows you couldn't resolve.
- Book the next session and rotate who picks the tickets.
After the session (facilitator, within 3 days)
- Publish the updated rubric with a one-line changelog.
- Add the calibrated tickets to the reference-example library.
- Share the disagreement table and the agreement rate with the team — including anyone who couldn't attend.
A healthy trend: agreement rate climbs and the disagreement table gets shorter over time. If it doesn't, your rubric is underspecified or your tickets are too easy — pick harder ones next round.
Continue exploring