Sampling math for small QA teams
Small teams review a random slice of tickets and report it like it's precise. The math says it's mostly noise — here's how many you actually need instead.

The number that won't sit still
Every month you pull twenty tickets at random, score them against the rubric, and report a team quality number to whoever asks. And every month it moves — 91 in March, 85 in April, 89 in May — even though nothing about how the team works has changed. You just looked at a different twenty tickets.
That isn't something you fix by being stricter with the rubric. It's a sampling problem, and for small teams it's nearly baked into the method most of us inherited.
Two percent of nothing
Start with how little manual QA actually covers. Across the industry, review reaches roughly this share of all support interactions:
For a large contact centre, 2% of a firehose is still a workable sample. For a five-person team handling a few hundred tickets a month, 2% is a handful of conversations — the kind of number a single strange ticket can swing on its own.
And most of what you pull back is fine. The industry-average Internal Quality Score sits high:
When the typical conversation already passes, a random sample spends most of its budget confirming that good agents did good work — and only now and then lands on the reopened, mishandled, quietly-churned ticket you actually needed to see.
The math nobody runs
Here is the part that gets skipped. A sample's precision scales with the square root of its size, not with the percentage of your volume it represents. To halve your margin of error you have to quadruple the tickets you read — punishing at the low end, where small teams live, and forgiving at the high end, where they don't. Review five tickets and see one fail, and the true fail rate consistent with that result runs from near-zero to well past half. That isn't a measurement; it's a shrug with a decimal point.
COPC ran the same math for CSAT, and it carries straight over to QA. An agent genuinely performing at 83.3% has this chance of randomly scoring 72% or worse in a month, purely from which conversations happened to get surveyed:
With the sample a small team can afford, a precise quality score isn't a measurement — it's a coin flip wearing a decimal point.
That comes off the volume a typical agent accumulates in a month —
— roughly the same order as the tickets a small team can hand-score per person. So when Priya's score drops eight points, you genuinely cannot tell whether Priya got worse or you drew a bad hand. Coaching to that number isn't rigour; it's reading tea leaves with a spreadsheet.
Random was always a compromise
Random sampling earns its keep on one axis: it's fair, and easy to defend when an agent asks why they got picked. But fairness to the scorecard isn't the same as finding what's broken — and the field has clocked the difference. Teams selecting review conversations purely at random are now the minority:
Down from 62% the year before. The centre of gravity is moving toward targeted, risk-based sampling, and for small teams that shift matters more, not less, because you have so few reviews to spend.
What to review when you can't review much
QA runs on a shoestring of reviews per agent nearly everywhere —
— so spend yours where the signal is.
Sample to a question, not a quota. For a defensible team-level read, work out the count you need for the margin of error you'll accept and stop chasing a percentage; the number is smaller than people fear, and because precision tracks the raw count, it barely grows as your volume does. For an individual, one month of random tickets won't get you there — pool several months, or over-sample deliberately.
Stratify instead of randomise. Pull the tickets most likely to hold a lesson: reopened threads, escalations, transfers, one-star CSAT, unusually long handle times, refunds, anything a customer had to chase twice. Same small number of reviews, aimed where problems actually live instead of the fat, boring middle.
Stop chasing a precise team score. With the sample a small team can afford, a two- or three-point monthly move is almost always noise, and treating it as signal will have you "fixing" things that never broke. Report trend and pattern, not decimals.
Keep your two jobs apart. Calibration — do reviewers agree on what "good" looks like — needs a shared handful of tickets everyone scores blind. Coaching needs specific, recent examples in front of the person who owns them. Neither needs a big random draw.
None of this makes QA softer. It makes it honest about what a small sample can and can't tell you — which is the whole point of running the math before you trust the number. If you're still settling what "good" even means, start with the QA benchmarks worth calibrating against and our glossary entry on Internal Quality Score.