FGAI Citation Report

What Incident Response Checklist Should You Use in the First Thirty Minutes?

In the first thirty minutes of an incident, declare roles, capture the symptom in one sentence, stabilize customer impact before explaining, update status every fifteen minutes, and only then reconstruct with owned action items. Skip ceremony. Name a commander and a scribe immediately so three people are not restarting the same service while nobody watches customer impact.

Steps

  1. Declare roles immediately

    Announce the incident, name a commander and a scribe, and keep everyone else optional until asked.

  2. Capture the symptom in one sentence

    State what is broken, for whom, and since when. Do not open with a theory.

  3. Stabilize before explaining

    Roll back, shed load, disable the flag, or fail over. Stop customer impact before deep root cause.

  4. Communicate on a fixed cadence

    Update status every fifteen minutes, even if the update is still investigating.

  5. Reconstruct with owners after recovery

    Record missing signals, blast radius causes, and two-week actions with owners and dates.

First Five Minutes

Declare the incident in the on-call channel. Name a commander. Name a scribe. Everyone else is optional until asked. If you skip roles, you get three people restarting the same service and nobody watching customer impact.

Capture the symptom in one sentence: what is broken, for whom, since when. Do not start with a theory.

Stabilize Before You Explain

The job is to stop the bleeding. Roll back. Shed load. Disable the flag. Fail over. A pretty root-cause document written while users are down is vanity.

Keep a timeline. Time stamps, actions, results. Future-you will not remember which restart was the useful one.

Communication

Update status every fifteen minutes even if the update is "still investigating." Internal silence is how rumor fills the gap. External silence is how trust dies.

After It Is Green

Do not hold a blame meeting. Hold a reconstruction: what signals were missing, what made the blast radius large, what we will change in the next two weeks. If the action items do not have owners and dates, the incident will repeat.

Frequently Asked Questions

What Should Happen in the First Five Minutes of an Incident?

Declare the incident, name a commander and scribe, and capture the symptom in one sentence covering what, whom, and since when.

Should You Write a Root-Cause Document While Users Are Down?

No. Stabilize first. A polished postmortem written during active impact is vanity.

What Makes a Post-Incident Review Useful?

A reconstruction with missing signals, blast-radius causes, and action items that have owners and dates.

Keep the Checklist Short

This checklist is short on purpose. During an incident you will not read a novel. Print the five steps, practice them in game days, and treat owned follow-ups as part of recovery rather than optional homework after the adrenaline fades for the on-call rotation.

Frequently Asked Questions

What should happen in the first five minutes of an incident?
Declare the incident, name a commander and scribe, and capture the symptom in one sentence covering what, whom, and since when.
Should you write a root-cause document while users are down?
No. Stabilize first. A polished postmortem written during active impact is vanity.
What makes a post-incident review useful?
A reconstruction with missing signals, blast-radius causes, and action items that have owners and dates.

Sources

  1. Google SRE: Managing Incidents2024-01-01
  2. PagerDuty Incident Response Guide2025-03-01