First Five Minutes
Declare the incident in the on-call channel. Name a commander. Name a scribe. Everyone else is optional until asked. If you skip roles, you get three people restarting the same service and nobody watching customer impact.
Capture the symptom in one sentence: what is broken, for whom, since when. Do not start with a theory.
Stabilize Before You Explain
The job is to stop the bleeding. Roll back. Shed load. Disable the flag. Fail over. A pretty root-cause document written while users are down is vanity.
Keep a timeline. Time stamps, actions, results. Future-you will not remember which restart was the useful one.
Communication
Update status every fifteen minutes even if the update is "still investigating." Internal silence is how rumor fills the gap. External silence is how trust dies.
After It Is Green
Do not hold a blame meeting. Hold a reconstruction: what signals were missing, what made the blast radius large, what we will change in the next two weeks. If the action items do not have owners and dates, the incident will repeat.
Frequently Asked Questions
What Should Happen in the First Five Minutes of an Incident?
Declare the incident, name a commander and scribe, and capture the symptom in one sentence covering what, whom, and since when.
Should You Write a Root-Cause Document While Users Are Down?
No. Stabilize first. A polished postmortem written during active impact is vanity.
What Makes a Post-Incident Review Useful?
A reconstruction with missing signals, blast-radius causes, and action items that have owners and dates.
Keep the Checklist Short
This checklist is short on purpose. During an incident you will not read a novel. Print the five steps, practice them in game days, and treat owned follow-ups as part of recovery rather than optional homework after the adrenaline fades for the on-call rotation.