Calibration: why two reviewers disagree on the same ticket
Hand the same ticket to two reviewers and you'll often get two very different scores. Calibration is the unglamorous fix, and the spread between reviewers is the metric that matters.

Two reviewers open the same ticket. One scores it 92; the other scores it 68. Same conversation, same scorecard, same customer, and a 24-point gap that lands on the rep's monthly review like a verdict decided by a coin toss.
That gap is not a rounding error. It is the everyday reality of quality assurance run without calibration, and it quietly poisons everything downstream: the coaching conversation, the bonus decision, and the rep's trust that any of it means something.
Disagreement is built in before a reviewer opens the ticket
Before you blame the reviewers, look at what you handed them. The typical scorecard is not short:
Every one of those categories is another place two honest people can read the same sentence differently. Some are close to mechanical, did the agent solve the problem, was the account verified, and they score high on agreement because there is a fact of the matter. The trouble is the soft categories, and teams weight them very unevenly:
That split is the whole story. Reviewers rarely fight over whether a refund was issued. They fight over whether "Thanks, sorted for you" reads as empathy or as curtness, and the more subjective categories you stack onto a scorecard, the more surface area you create for two totals to diverge. Not because either reviewer is wrong, but because the scorecard never told them what "good" looks like.
The scale you pick decides how much you'll argue
Rating-scale design is the quiet lever. The most common QA scale is the bluntest one available:
A pass/fail scale is crude, but it is hard to disagree about: the work either cleared the bar or it didn't. Widen the scale to five or ten points and you hand each reviewer a private ruler. One person's 7 is another's 9, and neither can quite say why.
This is exactly what calibration sessions exist to fix. You take one ticket, have several reviewers score it independently, then lay the scores side by side and argue until the definitions tighten. The output is not a "correct" score. It is a shared understanding of where the line sits, written back into the scorecard as worked examples. Do it monthly and drift stays small. Skip it for a quarter and every reviewer quietly reinvents the standard inside their own head.
A QA score isn't a measurement you take off the conversation. It's an agreement you build between the people reading it.
Thin samples make every disagreement expensive
You could tolerate some reviewer drift if reviews were plentiful. They are not:
A single review often stands in for hundreds of conversations no one will ever open. When coverage is that thin, one reviewer's idiosyncrasy isn't noise that averages out, it is the entire signal for that rep this month. And the number everyone chases rests on the same foundation:
That benchmark gets quoted in board decks as if it were read off a thermometer. But an IQS is only ever the sum of human judgements against a scorecard. If your reviewers have not calibrated, your 88% and someone else's 88% are not the same measurement. They may not even be the same measurement inside your own team.
What good looks like
Calibration is unglamorous, and it works. The practical version: pull one real ticket a week, have every reviewer score it blind, then meet for twenty minutes to reconcile. Track the spread between your highest and lowest scorer over time, that spread, not the average score, is the real health metric of a QA program. When it shrinks, your scores begin to mean something. When any two reviewers can land within a few points of each other, the coaching conversation finally stops being an argument about the referee.
Until then, treat every individual QA score as a range, not a verdict. The disagreement isn't a flaw in your people; it is information about your scorecard, and it is the cheapest feedback your quality program will ever get. If you want fewer of these fights, start with the instrument: here is how to build a support scorecard reviewers can actually agree on, and how quality practices shift as teams grow.