Incident & status-page comms pack
Copy-paste status-page and customer update templates for every stage of an outage — acknowledgement through post-incident follow-up — so your team can post clear, calm updates while the incident is still hot.
Use one block per stage, in order. Copy it, replace every [placeholder], and post. The golden rule: every update ends with a promised next-update time, and you never miss it — silence is what turns an outage into a trust problem.
Severity sets the cadence
Pick the severity first — it decides how often you owe customers an update and who signs off the wording.
| Severity | Customer impact | Post an update every | Who approves the post |
|---|---|---|---|
| SEV1 | Full outage, or data/security at risk | 30 min | Incident commander |
| SEV2 | Major feature down; workaround exists | 60 min | Incident commander |
| SEV3 | Degraded or slow; most customers unaffected | At milestones only | On-call lead |
Placeholders used across the blocks
| Placeholder | Fill with |
|---|---|
[TIME UTC] | Timestamp of this update, in UTC (e.g. 14:05 UTC) |
[SERVICE] | The product or component the customer recognises |
[IMPACT] | What a customer literally can't do, in their words |
[WORKAROUND] | A temporary way around it, or "No workaround yet." |
[NEXT UPDATE] | The UTC time you promise the next post by |
[SUPPORT CHANNEL] | Where to reach a human (email, chat, portal URL) |
Before you post (run this on every update)
- Written in plain customer language — no internal codenames, service IDs, or ticket numbers
- States impact from the customer's side ("you can't check out"), not the system's ("the checkout consumer lagged")
- Every timestamp carries a timezone — default to UTC
- Ends with a promised next-update time
- No blame, no individual names, no ETA you can't stand behind
- For SEV1/SEV2, approved by the incident commander before it goes live
1. Initial acknowledgement — status: Investigating
Post within the first few minutes of confirming impact. Acknowledge before you understand the cause; speed matters more than detail here.
[TIME UTC] — Investigating
We're aware of an issue affecting [SERVICE] and are investigating.
Customers may be [IMPACT, e.g. "unable to log in" / "seeing slow load times"].
We don't yet have a cause or an ETA. Next update by [NEXT UPDATE].
2. Investigating update — status: Investigating
Send this the moment the cadence clock runs out, even if nothing has changed. "Still working on it" is a valid, trust-keeping update — a skipped update is not.
[TIME UTC] — Investigating (update)
We're still investigating the issue with [SERVICE].
[WHAT WE'VE RULED OUT, or WHAT WE'RE SEEING].
Impact is unchanged: [IMPACT]. [WORKAROUND]
Next update by [NEXT UPDATE].
3. Identified / Monitoring
Move to Identified only once you know the cause; move to Monitoring only after the fix is live and the metrics are actually recovering. Don't skip ahead to reassure people.
Identified — status: Identified
[TIME UTC] — Identified
We've identified the cause of the issue affecting [SERVICE]:
[PLAIN-LANGUAGE CAUSE — no internal jargon].
A fix is [in progress / being deployed]. Impact: [IMPACT].
Estimated recovery: [ETA, or "still assessing"]. Next update by [NEXT UPDATE].
Monitoring — status: Monitoring
[TIME UTC] — Monitoring
We've deployed a fix for [SERVICE] and are monitoring recovery.
[SIGNAL, e.g. "Error rates are returning to normal."]
Some customers may still see [RESIDUAL IMPACT, or "no remaining impact"].
Next update by [NEXT UPDATE].
4. Resolved — status: Resolved
Only mark resolved once the metrics have held steady for a set window (e.g. 15–30 min). A premature "resolved" you have to reopen costs more trust than the outage did.
[TIME UTC] — Resolved
The issue affecting [SERVICE] is resolved as of [TIME UTC].
Service has been stable for [DURATION, e.g. "the last 30 minutes"].
[ONE-LINE CAUSE, if you can share it.]
If you're still seeing problems, contact [SUPPORT CHANNEL].
A short post-incident summary will follow within [TIMEFRAME, e.g. "3 business days"].
Thanks for your patience.
5. Post-incident follow-up — sent after resolution
Send to affected customers within the timeframe you promised. Explain impact and prevention in plain language: no jargon, no blame, no over-promising. This is where you earn the trust back.
Subject: What happened with [SERVICE] on [DATE]
Hi [NAME / there],
On [DATE], between [START TIME UTC] and [END TIME UTC], [SERVICE] was [IMPACT].
This affected [WHO / roughly how many customers]. We're sorry for the disruption.
What happened
[Plain-language cause, 2–3 sentences. No internal component names.]
What we did
[Detection → mitigation → resolution, briefly and in order.]
What we're changing so it doesn't recur
- [Prevention item 1]
- [Prevention item 2]
[If SLA credits apply: how affected customers can claim them.]
Questions? Reply to this email or reach us at [SUPPORT CHANNEL].
— [NAME], [ROLE]
After every incident (close-out checklist)
- Status page set back to Operational for all affected components
- Timeline captured with timestamps: detected, first post, cause identified, mitigated, resolved
- Customer-facing follow-up sent within the committed timeframe
- Blameless post-mortem scheduled with everyone who was on the call
- Action items logged with a named owner and a due date each
- Support macros and canned replies updated in case the issue recurs
Continue exploring