The QA rubric that predicts customer outcomes
Most QA scorecards reward tidy tickets and miss the customers who churn. Here is how to weight a rubric toward resolution, accuracy and tone — the things that predict outcomes.

A rep on my old team once posted a 98% on a QA review. Every box on the scorecard was ticked: warm greeting, correct branding, tidy tags, clean sign-off. The customer on the other end had already churned two weeks earlier. The rubric never noticed, because the rubric was never measuring whether the customer got what they came for.
This is the quiet failure of most quality programs. We build scorecards that reward neatness and punish typos, then act surprised when a team full of high scorers still bleeds retention. If your rubric can hand out a 98% on a ticket that ended in a cancellation, your rubric is measuring the wrong thing.
Compliance is easy to score and easy to ignore
The average support team evaluates a sprawling list of categories per conversation — far more than anyone can weight meaningfully.
Most of those categories are compliance checks: did the agent use the template, apply the tag, follow the greeting script. They're easy to score because they're binary and observable. They're also close to worthless as predictors, because no customer has ever renewed a contract on the strength of a correctly applied macro.
Here's the uncomfortable evidence. When one QA platform analysed a very large ticket sample, internal quality scores and customer satisfaction simply did not move together.
Read that carefully. It doesn't mean QA is useless. It means a typical QA score and a CSAT score are measuring two different things — and if your scorecard is built mostly from compliance items, you've built an instrument that's precise about the stuff customers never feel.
A rubric that scores everything weights nothing.
Weight for resolution, accuracy, and tone
The fix isn't a longer scorecard. It's a shorter one, weighted hard toward the three things that actually track with outcomes: did the customer's problem get solved, was the information correct, and did the exchange feel human.
Resolution is the heavyweight. It's already the most-scored category in the business — but empathy lags badly behind.
That gap is the whole problem in one line. Teams check whether the issue was solved far more often than whether the customer was treated like a person, even though both drive loyalty. And resolution genuinely correlates with the outcome you care about — every point of first-contact resolution tends to move satisfaction by roughly the same amount:
So a rubric that predicts outcomes looks lopsided on purpose. Resolution and accuracy might carry half the total weight. Tone carries a meaningful chunk. Compliance items — tagging, formatting, template use — survive only as pass/fail hygiene flags that never inflate a score. They don't earn points; they just catch sloppiness.
Make the scale force a judgement
How you score matters as much as what you score. The most common rating scale in the industry is a blunt one.
A two-point pass/fail scale is fine for hygiene items — a tag is either right or it isn't. But resolution and tone live on a spectrum, and collapsing them to pass/fail throws away the exact signal you're trying to capture. Score the outcome categories on a 3- or 4-point scale so a reviewer has to distinguish "technically resolved" from "resolved in a way the customer trusts."
And be honest about what the rubric is for. It exists because you can't coach from CSAT alone — the sample per agent is simply too thin to be reliable.
That's the real case for a good rubric. CSAT tells you the weather; a well-weighted QA score tells you which behaviours to change on Monday. The scorecard earns the right to be your coaching instrument precisely because it's stable enough to act on, where an individual's monthly survey scores are mostly noise.
Build the rubric backwards
Start from the outcome, not the checklist. List the behaviours you can actually observe in a transcript that plausibly move resolution, accuracy, and trust. Weight those. Demote everything else to a hygiene flag. Then sample enough conversations that the score means something, and carry the feedback into your 1:1s rather than filing it in a dashboard nobody opens.
Pick the few things that predict whether the customer comes back, and let the rest fall away. If you want to sanity-check your weightings against where the industry actually sits, the benchmarks are a decent mirror — and if the line between an internal quality score and CSAT still feels fuzzy, the glossary untangles them.