Loading...

About the Incident Checklist Editor

A checklist exists because incidents degrade judgement. Under time pressure people skip the step that feels obvious, forget to tell anyone, or start fixing before anyone has confirmed what is broken. A written sequence is not a substitute for expertise — it is what stops experienced people missing the routine parts while concentrating on the hard one.

The steps that get skipped are almost always communication and role assignment, not diagnosis. Naming an incident commander, opening a channel, and posting an initial status update take two minutes and are the difference between a coordinated response and five people independently investigating the same thing.

Keep it short enough to actually follow. A checklist that runs to forty items will be abandoned at item six during a real incident. Cover declaration, roles, communication, mitigation and stand-down, and leave the diagnosis to the people doing it — a checklist that tries to encode troubleshooting becomes stale and misleading.

Frequently asked questions

What belongs in an incident checklist?

Declaration and severity, who is incident commander, where coordination happens, who tells customers and when, the mitigation attempt, and the stand-down with a follow-up owner. Deliberately not troubleshooting steps — those belong in per-service runbooks, and mixing them makes the checklist long enough that nobody reads it.

Who should be incident commander?

Whoever declared the incident, until they hand over explicitly. The role coordinates rather than fixes, and the most common failure is the commander also being the person deepest in the debugging — at which point nobody is tracking time, communications or whether the current approach is working. Handing over is normal and should be easy.

When should we tell customers?

Earlier than feels comfortable, and before you know the cause. An early acknowledgement that you are aware and investigating costs nothing and buys patience; silence while you find the answer reads as not noticing. Agree the threshold in advance so it is not a judgement call at three in the morning.

How does this relate to a postmortem?

The checklist should end by assigning someone to write one and setting a date. Capture the timeline during the incident rather than reconstructing it afterwards — timestamps of when things were noticed, attempted and changed are far harder to recover later, and they are the most valuable part of the document.

Should severity levels be in the checklist?

Yes, with concrete criteria rather than adjectives. "Customer-facing and no workaround" is actionable; "major impact" invites debate during the worst possible moment. Two or three levels with clear thresholds work better than five that people argue about while the service is down.

Need this managed for you, not just automated?

We're also a hands-on DevOps consultancy — Kubernetes, CI/CD, and cloud infrastructure.

Explore Our Services