Your CSAT went up and your customers trust you less
AI can lift every surface metric while quietly eroding the thing underneath them. A warning about measurement illusions — and how to see through them.

Here's a scenario that should worry you more than a bad quarter: every support metric on your dashboard is up. Response time down, CSAT up, handle time down, deflection up. The board is pleased. And your customers trust you less than they did a year ago. This is not a paradox — it's the predictable result of measuring the surface while automation reshapes what's underneath. It's also one of the easiest traps to fall into with AI.
Start with the uncomfortable data, because it sits directly against the story most dashboards tell.
Most customers don't want AI in their support experience, and a majority would consider leaving over it. You can hold those facts in one hand and a rising CSAT score in the other — because they measure different things, and the gap between them is exactly where trust leaks out.
Why the surface metrics rise
AI is genuinely good at the things CSAT and speed metrics reward. It responds instantly. It's unfailingly polite. It resolves the easy cases cleanly — and the customers with easy cases are happy, and they're also the ones most likely to answer a CSAT survey, because their experience was quick and painless. Your response pool skews toward the people the machine served well.
Meanwhile, the customer with the genuinely hard problem — the one the bot looped for ten minutes before a human finally appeared — is not filling out your survey. They've escalated into a channel you measure differently, or they've given up.
CSAT measures the customers who stayed to answer. Trust is measured by the ones who quietly left.
That's the core illusion. CSAT is a survey of the people who stayed to take it. It structurally under-samples the customers automation failed — which means the more customers automation quietly fails, the cleaner your CSAT can look. The metric improves as the underlying experience polarises.
What actually erodes
The hard cases get harder. As AI absorbs the routine, the human queue concentrates into the difficult and the emotional. If you haven't resourced for that (see the org chart after AI), your best-remembered interactions — the high-stakes ones — get worse even as your average score climbs.
The friction is invisible to your instruments. The doom-loop, the un-findable answer, the "did that resolve your issue?" that the customer clicked yes on just to escape — none of these fire an alarm on a standard dashboard. They surface later, somewhere your support metrics don't reach.
Trust is a lagging indicator, and it's brittle. People forgive a lot, but not repeatedly, and not when they feel deflected on purpose.
That PwC figure is old — treat it as the classic baseline it is — but the direction has only intensified since. One bad, automated-feeling experience is enough for a meaningful share of customers to leave, and they rarely tell you why. They just don't come back.
How to measure the thing that's actually eroding
You cannot fix what your instruments can't see. So change the instruments:
- Segment CSAT by path. Score customers served by AI separately from those who reached a human, and separately again for those who did both. If the AI-path score is high but that cohort's retention is low, you've found the leak.
- Watch repeat contacts and reopens, not just first-contact resolution. A resolved-then-reopened ticket is a trust event, not a stat.
- Instrument the failures, not just the successes. Sample the abandoned chats, the escalations, the "no" responses. The customers who didn't stay to be surveyed have the most to tell you.
- Track a cohort's behaviour, not a survey's mood. Do customers routed through automation buy again, renew, refer? Behaviour is trust; a survey is a feeling captured from a biased sample.
None of this makes the dashboard prettier. It makes it true, which is a different and more useful thing. The goal was never a high CSAT — it was customers who stay. When the two diverge, and AI is very good at making them diverge, believe the retention number, not the survey. And whatever you do, don't let a rising score talk you out of noticing the people who quietly left. They're the report that matters. (For the deflection-rate version of this same illusion, see deflection is a vanity metric.)