Calibration clinic: ten disputed support interactions
A practical QA clinic for separating preference from risk, scoring evidence consistently, and turning disagreement into a better system.

Calibration often becomes a room full of experienced people defending different tastes. One person dislikes the greeting. Another thinks the reply is too long. A third gives full marks because the issue was solved. None of that creates a reliable quality system.
Traditional manual QA may inspect only
of conversations. When the sample is that small, every scored interaction carries disproportionate weight. Calibration must therefore test the definition behind the score, not pressure reviewers into matching a senior person's number.
The clinic method
Choose ten interactions that contain genuine ambiguity: an approved exception, a partial outage, a vulnerable customer, a policy conflict, an AI-assisted reply, a transfer, a technically correct but confusing answer, a safe refusal, an unnecessary escalation, and a resolved case with poor ownership.
Each reviewer scores independently before discussion. For every failed item, they must cite the observable evidence and the written standard. “It felt off” is a useful signal for discussion, not a score rationale.
Ten disputes worth practicing
- Warm but incomplete: the agent is empathetic and misses one required security step.
- Cold but complete: every fact is correct; the customer cannot easily find the answer.
- The sensible exception: policy says no, but documented discretion permits a narrow remedy.
- The escalated non-answer: the agent transfers quickly without trying the work they own.
- The heroic workaround: the customer is helped, but the action creates an audit or privacy risk.
- The AI draft: fluent and mostly accurate, with one unsupported claim the agent failed to verify.
- The repeat customer: the reply solves today's message while ignoring the unresolved original goal.
- The long explanation: accurate context is included, but the action is buried below it.
- The safe refusal: the agent protects a boundary and gives a viable alternative.
- The broken system: the agent follows the process exactly and the process produces harm.
Resolve disagreement in the right order
First ask whether reviewers saw the same evidence. Then ask whether the standard covers the situation. Only then ask whether the scoring scale is clear. Do not rewrite a scorecard when the real problem is missing context or an unmade policy decision.
Use three labels for disputed items:
- Reviewer drift: the standard is clear; interpretation has moved.
- Standard gap: reasonable reviewers cannot reach the same answer from the written rule.
- System defect: the interaction exposes broken policy, knowledge, tooling, or authority.