When production breaks, support pays the bill
Engineering gets a war room, an incident commander, and a retro. Support gets a tripled queue and a macro written at 2 a.m. The support incident runbook that fixes the asymmetry.

The deploy goes red at 14:07. By 14:12 engineering has a war room, an incident commander, and a timeline document. By 14:27 the support queue has tripled, the chat widget is a wall of the same question asked four hundred ways, and the agents answering it are learning about the outage from the customers. Nobody invited support to the war room. Somebody should have — because most of the incident's bill is about to land on support's side of the house.
The bill is bigger than the invoice
Infrastructure people measure outage costs carefully, and their numbers are sobering on their own:
But read what those tallies count: lost revenue, recovery labour, SLA penalties on the infrastructure side. The support-side costs — the contact spike, the overtime, the breached response SLAs, the surveys that arrive the following week — hide in a different budget line and rarely make the incident's price tag. And the deferred cost is the largest of all. An outage is a machine for manufacturing bad experiences in bulk:
'Multiple bad experiences' usually accumulate one customer at a time, over months. An incident delivers them to your entire customer base in a single afternoon — the failed action, the confusing error, the unanswered chat, the second attempt that also failed. Four bad experiences before dinner. The switching math doesn't care that they shared a root cause.
Incidents are process failures — which is good news
The infrastructure research holds a second lesson support should steal:
Overwhelmingly, things break because procedures weren't followed or didn't exist — not because the universe is hostile. The same is true of support's incident response. The tripled queue isn't the failure; the failure is meeting it with improvisation: the macro drafted mid-spike, the status page nobody owns, the agent guessing at an ETA because nobody told them anything. All of that is process you can write before you need it.
Engineering gets a war room, a commander, and a retro. Support gets a tripled queue and a macro written at 2 a.m. That asymmetry is a choice.
The support incident runbook
- A seat in the war room. One support person in the incident channel, with a speaking part: they carry what customers are actually seeing (which is often diagnostic — support frequently knows the blast radius before monitoring does) and carry back what agents can say. Support learning about the outage from customers is the single most fixable failure on this list.
- Treat the status page as deflection infrastructure. Its job is to absorb the queue spike before it happens. That only works if it updates before the contact wave — specific, timestamped, honest about what's known and when the next update comes. A status page that says 'all systems operational' twenty minutes into an incident doesn't just fail to deflect; it teaches customers to never check it again.
- Pre-write the skeletons. The acknowledgement macro, the workaround template, the resolution notice, the apology — drafted calm, at leisure, with legal's blessing, with blanks for the specifics. The 2 a.m. version written mid-incident is always worse. Ours are in the incident comms pack, free to steal.
- Set ETAs like they're promises, because they are. Customer patience is already thinner every year —
— and mid-incident it thins by the minute. Publish update cadence ('next update at :30') rather than resolution guesses; a missed ETA during an incident is a second incident.
- Tag everything. Every incident-related ticket gets the incident's tag, in the moment. After resolution, support owns the customer-impact evidence — how many affected, which segments, what they tried, what they threatened. That corpus is support's ticket into the postmortem, where 'customer impact' should be a section support writes, not a sentence engineering estimates.
- Watch the aftershock. The queue doesn't end when the incident does: reopens, billing questions, trust wobbles, and the backlog you built during the spike — which needs the aging-tail discipline from backlog triage, not a week of silent grinding. And decide your SLA-pause policy now, in daylight — SLAs that protect customers covers why a breached-but-honest clock beats a paused-and-quiet one.
An incident is the one day your whole customer base experiences your support at once. Engineering rehearses for that day. The queue deserves the same.