Concurrency: the hidden cap on chat quality
The three-times-cheaper math behind live chat lives or dies on one dial nobody watches: how many conversations an agent runs at once. Here is what one, two and three concurrent chats actually cost.

Every workforce planner has run the same seductive sum. A phone agent handles one conversation at a time; a chat agent can hold two or three at once; therefore chat costs a third of what voice does per contact. That multiplier is the entire financial case for live chat, and it gets baked straight into the staffing model. What almost nobody writes down is that concurrency — the number of conversations one agent runs in parallel — is not a free lever. It is a quality dial being turned with no readout attached.
Holding three chats is not the same as working three times faster. It is switching between three tasks, and switching has a price. Cognitive-psychology research finds that shuffling between tasks can cost as much as
. The support version is mundane but relentless: you read a message, load the customer's history, begin a reply, an alert yanks you to a second thread, and when you return you reread your own half-finished sentence to remember where you were. The same tax that plagues tool-hopping — the subject of the toggling tax — operates inside a single chat console, one window at a time.
That tax lands on the one thing chat is bought for: speed. Customers now get a first reply on live chat in roughly
, and that expectation is the channel's whole reputation. Concurrency is what puts it at risk. The customer in your third window is not experiencing a fast channel; they are watching a typing indicator stall while you close out the other two. When the wait stretches, people rarely complain — they leave. LiveChat's report puts queue abandonment at
, and every additional parallel thread quietly lengthens the wait that produces it.
Here is the shape of the trade-off. Going from one chat to two is nearly free: agents fill the dead air while a customer types or reads, and throughput almost doubles. Two to three is where the curve bends — you buy a little capacity by spending a lot of quality, because the agent is now holding three contexts and arbitrating whose reply goes first. Past three, you are running a queue inside each agent's head, and handle time on every open conversation rises to pay for it.
The second chat is almost free. The third is where you buy capacity with quality — and the fourth is a queue you've hidden inside one agent's attention.
The satisfaction numbers follow the same bend. Chat's folklore says it is the happy channel, yet measured across rated conversations its CSAT averages
— a fair bit below the legend. Comm100's benchmark, drawn from 220 million interactions across 18 industries, puts overall live-chat satisfaction at
: steady, but hardly the runaway winner the staffing sum assumes. Neither figure improves when you push agents to their concurrency ceiling. The switching cost above guarantees the opposite, so raising the dial doesn't just lengthen replies — it shaves the CSAT benchmark you report.
Ask the people who actually run chat floors and the answer is consistent. Practitioners converge on
concurrent chats as the working ceiling, with a firm "never more than three" and new agents held at one until they are fluent. The number moves with complexity: simple order-status chats tolerate more parallelism; anything technical or emotional needs an agent's whole attention, and stacking those is how the wrong answer gets sent to the wrong customer.
The deeper problem is that concurrency is invisible on most dashboards. You see AHT, you see CSAT, you see occupancy — but not the concurrency that moved all three, so you tune the metrics you can see and never touch the lever beneath them. The fix is cheap: log the number of open chats at the moment each reply is sent, then plot response time and CSAT against it. There is almost always a knee — the point where one more chat stops adding capacity and starts subtracting quality. Set the cap just below it, treat the three-times-cheaper multiplier as a hypothesis to test per team rather than a constant to assume, and staff to the concurrency your CSAT can actually survive.