The operating view
What this topic covers
For QA analysts, team leads, support managers, and operations leaders creating a program or replacing a punitive scorecard with useful evidence.
Quality assurance answers a harder question than whether a ticket closed: was the help accurate, complete, safe, understandable, appropriately human, and likely to hold? A useful program turns those standards into observable behavior and then converts review findings into coaching, knowledge, process, policy, and product changes.
QA loses trust when the rubric rewards style over resolution, reviewers disagree without examining why, or scores become a surprise performance weapon. The score is not the product. The product is shared judgment: agents know what good looks like, leaders know where the system fails, and customers receive more dependable outcomes.
Automation can expand coverage, but coverage is not validity. An automated evaluator needs the same clear rubric, representative test set, calibration, appeal path, drift monitoring, and governance as a human review process. Use machines to find patterns and prioritize attention; keep humans responsible for ambiguous, high-impact judgments.
Score observable evidence
A reviewer should be able to point to the conversation or required action supporting a judgment. Avoid criteria such as 'went above and beyond' unless the behavior is defined.
Separate severity from arithmetic
A critical privacy or accuracy failure may matter more than several minor style successes. Use critical-failure rules rather than letting weighted averages bury risk.
Treat disagreement as data
Calibration gaps reveal unclear standards, missing context, or legitimate judgment calls. Investigate them instead of forcing superficial consensus.
Core practice
Build a rubric from the customer promise
Translate policy and desired outcomes into a compact set of observable dimensions.
Start with accuracy, resolution, clarity, ownership, empathy appropriate to the situation, process compliance, and safety. Remove criteria that duplicate one another or measure personal writing preference. For each item, include examples of meets, misses, and not applicable.
Pilot the rubric on real conversations spanning channels, issue types, and complexity. If reviewers repeatedly need exceptions, revise the standard before attaching consequences. Version the rubric and explain changes so trend lines are not mistaken for performance movement.
Operator checks
- Criteria are observable
- Critical failures are explicit
- Not-applicable handling is defined
Core practice
Sample for the question you need to answer
Random coverage alone can miss rare, severe, and operationally important work.
Use a representative sample to estimate ordinary performance, then add targeted samples for new hires, new policies, escalations, low-confidence automation, high-risk intents, and changed workflows. Keep targeted findings distinct from population estimates so a risk queue does not make the whole team appear worse.
Small samples create noisy individual scores. Use them for coaching evidence and themes, not false precision. If a score will affect employment or compensation, the process needs sufficient evidence, review consistency, context, and an appeal path.
Operator checks
- Sample purpose is named
- Risk samples are labeled separately
- Consequences match evidence strength
Core practice
Calibrate judgment, not just scores
Use disagreement to improve the rubric and the shared interpretation behind it.
In calibration, reviewers score independently before discussing evidence. Compare dimensions, not only totals. Ask which fact changed the judgment, whether the rubric supplied enough guidance, and whether unseen context was assumed. Record the decision and add a worked example where it will prevent future ambiguity.
Calibrate on a cadence and whenever policy, channel, evaluator, or automation changes. Agreement percentage can hide systematic bias, so inspect which reviewer, team, language, or criterion drives differences.
Operator checks
- Reviewers score before discussion
- Decisions become examples
- Bias is checked by reviewer and segment
Core practice
Turn findings into one testable behavior
A review creates value only when the right person or team can act on it.
Coaching should connect evidence, customer impact, desired behavior, practice, and follow-up. Choose one or two changes the agent can control. If the failure came from a missing permission, contradictory article, broken integration, workload design, or policy, assign it to the system owner rather than disguising it as an individual coaching issue.
Measure whether the behavior changed in later work and whether the customer outcome improved. A completed coaching meeting is activity; durable improvement is the result.
Operator checks
- Feedback cites evidence
- System causes have system owners
- Follow-up checks behavior and outcome
Progression
QA maturity
Maturity grows as quality evidence becomes more trusted and more useful outside the scorecard.
- Stage 1
Inspection
A manager samples convenient tickets and gives informal feedback.
- Standards vary by reviewer
- Coverage and findings are unclear
Next move: Create a compact rubric, sampling plan, and review record.
- Stage 2
Calibrated
Reviewers use shared examples, regular calibration, and an appeal path.
- Disagreement is measured
- Critical failures are governed
Next move: Connect findings to coaching and system owners.
- Stage 3
Learning
QA themes change knowledge, workflow, policy, and training as well as agent behavior.
- System causes are separated
- Follow-up measures improvement
Next move: Add targeted risk coverage and validated automation.
- Stage 4
Predictive
Quality signals identify emerging customer and operational risk before headline outcomes move.
- Coverage is risk-aware
- Evaluators are monitored for drift
Next move: Keep human judgment, fairness, and customer outcomes central as scale grows.
Use the framework
Diagnose a quality finding
Before assigning coaching, locate the cause the organization can actually change.
- 01
What observable behavior or omission occurred?
Quote the evidence and required standard without interpreting intent.
Output · A specific finding. - 02
What customer or business risk did it create?
Separate style preference from accuracy, effort, trust, safety, or resolution impact.
Output · A severity and consequence. - 03
Was the agent able and equipped to succeed?
Check knowledge, permissions, tools, workload, policy, and training.
Output · An individual or system owner. - 04
How will improvement be verified?
Choose later work and an outcome signal rather than meeting completion.
Output · A follow-up date and evidence.
Common questions
Frequently asked questions
How many support conversations should QA review?
It depends on the question and consequence. Use enough representative work to identify team patterns, then targeted samples for risk and coaching. Do not claim precise individual performance from a handful of cases.
Should QA scores affect compensation?
Only with unusually strong safeguards: stable standards, sufficient evidence, calibration, context, transparency, and appeal. Programs earn more trust and improvement when scores begin as developmental evidence rather than punishment.
Can AI replace human QA reviewers?
It can expand screening and consistency for well-defined criteria, but humans remain necessary for calibration, ambiguity, fairness, policy interpretation, and high-impact decisions. Validate each criterion rather than treating one overall correlation as proof.
What should happen after calibration disagreement?
Record the evidence and decision, clarify the rubric, add an example, and check whether earlier scores need correction. Agreement achieved only inside the meeting is not a durable control.
Start with something useful
Curated reference shelf
QA scorecard and rubric
A complete starting point for dimensions, evidence, and critical failures.
Open resource TemplateCalibration agenda
A repeatable structure for aligning judgment with evidence.
Open resource BenchmarkQuality-score benchmarks
Sourced external context for internal quality scores.
Open resource TemplateCoaching 1:1
Connect QA evidence to practice and follow-up.
Open resource