Skip to content

Incident Response — Levels, Roles, Runbooks ​

Tại sao incident response cần cấu trúc ​

Một incident là stressful. People panic. Communication breaks down. Multiple teams try to fix different things at once. Communication cascades upward without clear info flow.

Result: incident takes 2 hours to fix when actual recovery is 15 minutes.

Structured incident response with clear levels, roles, and runbooks turns chaos into coordinated action:

  1. Levels — How severe is this? How many people needed?
  2. Roles — Who does what? Who decides? Who communicates?
  3. Runbooks — Steps to take, not "figure it out" under pressure

Internal Model: Incident severity levels ​

Level definitions ​

Most organizations use 3–4 levels:

Level 1 (Critical)

  • Customer-facing service completely down
  • Revenue impact or SLO breach imminent
  • Response time: < 5 minutes to engage
  • Team size: 10–20 people (full war room)

Example:

  • Primary database down (no API responses)
  • Authentication service returning 5xx (no logins possible)
  • Payment processing failing (transactions not completing)

Level 2 (Major)

  • Service degraded but not down
  • Customer experience affected (slow, errors)
  • SLO may be breached soon
  • Response time: < 30 minutes to engage
  • Team size: 5–10 people

Example:

  • API p99 latency doubled
  • 20% of requests returning errors
  • Search service slow but working

Level 3 (Minor)

  • Customer experience slightly affected
  • No SLO impact yet
  • Response time: < 2 hours
  • Team size: 1–3 people

Example:

  • Internal dashboard slow
  • Email alerting delayed by 5 minutes
  • Logging pipeline backlogged

Level 4 (Informational) (optional)

  • No customer impact
  • Just an anomaly to investigate
  • No page to on-call
  • Team size: async (best effort)

How to determine level ​

During incident triage, ask:

  1. Customer impact — how many users affected?
  2. Service degradation — is service down or slow?
  3. SLO impact — will this breach SLO?
  4. Business impact — revenue at risk?
If (customers affected AND service down) → Level 1
Else if (customers affected AND degraded) → Level 2
Else if (internal only) → Level 3
Else → Level 4

Level escalation ​

Incidents can escalate:

Level 3 investigation → "Oh, root cause is database issue"
                     → Now upstream database team involved
                     → Escalate to Level 2 (multiple teams needed)
                     
Level 2 incident → "Can't recover automatically"
                → "Need human decision from VP"
                → Escalate to Level 1 (executive engagement)

Rule: Level is determined by severity + scope, not by team. A widespread Level 3 can become Level 2 if it affects more users.

Internal Model: Roles during incident ​

Incident Commander (IC) ​

The single decision maker during incident. Not necessarily most senior person, but someone trained for this role.

Responsibilities:

  • Declare incident level and invoke appropriate response
  • Direct investigation ("SME, look at database logs")
  • Authorize changes ("Deploy the hotfix")
  • Decide escalation ("This needs VP approval")
  • Set timeline ("We have 30 minutes to stabilize before we need to escalate")

Why single IC?

  • Prevents multiple people issuing conflicting directions
  • Clear accountability (if decision is wrong, IC owns it)
  • Fast decision-making (no committee)

IC is NOT:

  • The person who fixes the problem
  • The person who knows the system best
  • The CEO in crisis

IC IS:

  • Someone who can coordinate teams
  • Someone who can say "stop, let's try a different approach"
  • Someone who can handle pressure

Subject Matter Expert (SME) / Incident Responder ​

People who actually diagnose and fix the issue.

For Level 1:

  • Multiple SMEs from different areas (database, API, infrastructure, etc.)
  • Each owns investigation of their component

For Level 2–3:

  • Usually 1–3 people with knowledge of the service

Key behavior:

  • SME reports status to IC frequently ("Found the issue, it's X")
  • SME asks IC before deploying changes ("IC, I have a fix, should I deploy?")
  • SME documents findings for postmortem

Communications Lead ​

Person who updates customers, status page, leadership.

Responsibilities:

  • Write initial notification ("We are investigating an issue affecting...")
  • Send periodic updates (every 15–30 minutes, or "status unchanged")
  • Final post-incident update ("Issue resolved at 2:45 PM. Root cause was...")
  • Answer customer inquiries

Why separate role?

  • SMEs focused on fixing, not on communication
  • Prevents SMEs from making conflicting public statements
  • Ensures consistent messaging

On-call lead ​

For Level 1 incidents, on-call lead ensures escalation and resource coordination.

Responsibilities:

  • Wake up the right people (page backend SMEs, database SMEs, etc.)
  • Ensure IC has authority to make decisions (sometimes VP approval needed)
  • Track whether incident is being resolved or escalating
  • Escalate to higher management if needed

Example incident command for Level 1 ​

13:45 - Issue detected
  ├─ Alert fired: "Database CPU at 95%"
  ├─ On-call engineer pages IC
  └─ IC acknowledges, enters war room

13:50 - Incident declared
  ├─ IC: "Level 1 declared. Database issue."
  ├─ IC pages: Backend SME, Database SME, Infrastructure SME
  └─ Communications lead starts status page: "Investigating issue"

13:55 - Investigation begins
  ├─ Database SME: "Queries running slow, 100K queries in queue"
  ├─ Backend SME: "Requests timing out waiting for DB"
  ├─ Infra SME: "No resource constraints, disk I/O not maxed"
  └─ IC: "Database query performance is bottleneck. Let's focus there."

14:00 - Root cause found
  ├─ Database SME: "Slow query is analytics job running unplanned"
  ├─ IC: "Kill the analytics job?"
  ├─ Database SME: "Yes, safe to kill."
  ├─ IC: "Do it."
  └─ Comms: "Issue root cause found, recovering"

14:05 - Recovery
  ├─ Analytics job killed
  ├─ Database CPU drops to 40%
  ├─ API response times return to normal
  ├─ SLO no longer breached
  └─ IC: "Incident recovered."

14:10 - Incident closed
  ├─ Comms: Final update posted
  ├─ All teams stand down
  ├─ IC: "Postmortem in 24 hours"

Key observations:

  • Single IC directed whole response
  • Clear roles (who investigates what)
  • Fast communication (status updated every 5 minutes)
  • Decision documented ("kill analytics job" authorized)

Internal Model: Runbooks ​

What is a runbook? ​

A runbook is a machine-readable sequence of steps to respond to a specific incident type.

NOT:

  • A long narrative explaining the system
  • An after-the-fact explanation
  • A training document

IS:

  • A checklist to follow during incident
  • Executable under pressure
  • Step by step with commands/queries to run

Runbook structure ​

Title: Database Slow Query Incident

Trigger: Database CPU > 90% for > 2 minutes

Level: 2 (escalate to Level 1 if can't recover in 30 min)

Steps:

1. Verify database is actually slow
   $ gcloud sql instances describe <INSTANCE>
   Look for: CPU_Utilization, current_connections
   
2. Check for long-running queries
   $ gcloud sql operations list --instance=<INSTANCE> --limit=20
   Filter: operations taking > 2 minutes
   
3. If found slow query:
   a. Identify the query (check logs)
   b. Determine if safe to kill
      - Is it a user-facing query? (check service name in logs)
      - Is it an analytical job? (check tags)
      - Decision: User-facing = don't kill. Analytics = safe to kill.
   c. Kill if safe
      $ gcloud sql operations cancel <OPERATION_ID>
      
4. If no slow query found:
   - Check disk I/O
   - Check memory
   - (If still stuck, escalate to Database team SME)
   
5. Verify recovery
   $ gcloud sql instances describe <INSTANCE>
   CPU should drop within 1 minute
   
6. If recovered, close incident
   If not recovered after 30 minutes, escalate to Level 1

Escalation:
- Contact: [email protected]
- Page: database-platform team

When to have a runbook ​

Not every possible incident needs a runbook. Prioritize:

  1. Incidents that happen > 2 times per year — worth documenting
  2. Incidents with time-sensitive steps — runbook prevents delay
  3. Incidents involving complex systems — runbook reduces errors

Don't create runbooks for:

  • Very rare incidents (< 1 per year)
  • Incidents requiring custom analysis
  • Incidents involving unclear root causes

Keeping runbooks up to date ​

Common problem: runbooks become stale.

Year 1: Runbook created, useful
Year 2: System architecture changed, but runbook not updated
Year 3: Runbook has wrong commands, on-call follows it and wastes time

Solution:

  • Every 6 months, review runbooks used in actual incidents
  • If runbook was used but contained outdated info, fix it immediately
  • Quarterly: SME audit to verify commands still work

Pro practice: Automated runbook testing.

Test: "Slow query runbook commands still work"
  - Monday morning at 2 AM, test the runbook
  - Run each command on non-production environment
  - Alert if any command fails
  - Fix failing commands

Incident roles on-call rotation ​

On-call for Level 1 incidents ​

Typically:

  • IC on-call: 1 person, rotates weekly/biweekly (most stressful)
  • Backend/Database/Infra SME on-call: 1 person per team, rotates
  • Communications on-call: Optional, can be rotated or assigned

Coverage:

  • 24/7 IC coverage (especially for critical services)
  • SME coverage aligned with when service gets traffic

On-call schedule ​

Week of June 24:
┌─────────────────────────────────────────────┐
│ IC: Alex                                    │
│ Backend SME: Bao                            │
│ Database SME: Carmen                        │
│ Infra SME: Diana                            │
│ Comms: Elena                                │
└─────────────────────────────────────────────┘

Each person on-call:
- Carries phone (or laptop if distributed)
- Responds to pages within 5 minutes
- Hands off to next person at week end

Pager rotation considerations ​

  • Prevent burnout: Max 1 week on-call per month
  • Distribute fairly: Track who was on-call when
  • Skill development: Senior engineer mentors junior during on-call
  • Compensation: Some organizations offer extra pay for on-call

Communication channels during incident ​

Incident war room setup ​

For Level 1, dedicated communication channel:

Slack channel: #incident-2025-06-24-db-issue

Timeline in channel:
  13:50 [IC:Alex]: Level 1 declared - database issue
  13:52 [DBE:Carmen]: DB CPU 95%, investigating query queue
  13:55 [Backend:Bao]: API p99 latency 5s (normal <100ms)
  13:58 [DBE:Carmen]: Found slow analytics job, safe to kill
  14:00 [IC:Alex]: Kill analytics job now
  14:02 [DBE:Carmen]: Killed. CPU dropping.
  14:05 [IC:Alex]: Recovered. Incident closed. Postmortem tomorrow.

Separate channels:

  • #incident-response: War room channel (for people actively responding)
  • #status-page: What to post publicly
  • #engineering-all: Executive summary (separate from real-time tactical details)

Common mistakes in incident response ​

Mistake 1: IC authority unclear ​

Multiple people make decisions:

  • Backend lead says "deploy hotfix"
  • Database lead says "wait, let's investigate first"
  • Result: contradictory actions, wasted time

Fix: Make IC authority crystal clear. Everyone defers to IC.

Mistake 2: Runbooks are narrative, not executable ​

Runbook says: "Check if database is slow"

But how? No commands provided. On-call has to figure out what tool to use. Wastes 10 minutes during incident.

Fix: Runbooks must have exact commands:

$ gcloud sql instances describe my-db --format=json | jq '.settings.backupConfiguration'

Mistake 3: No postmortem scheduled immediately ​

Incident ends at 2:45 PM. On-call is exhausted.

PM: "Let's have postmortem Friday."

Friday: People don't remember details. Postmortem is vague. Nothing learned.

Fix: Schedule postmortem immediately after incident ("Postmortem at 4 PM, 1 hour from now while memory fresh").

Mistake 4: Runbooks never tested ​

Runbook says: "Deploy hotfix X"

During incident: "hotfix X not in repo" or "build command is wrong"

On-call scrambles, wastes time.

Fix: Regularly test runbooks (at least quarterly).

Summary ​

Structured incident response enables fast recovery under pressure:

□ Levels define severity and response scope
□ IC provides single decision authority (prevents thrashing)
□ Roles distribute responsibilities (investigation, communication, execution)
□ Runbooks encode knowledge in executable steps
□ On-call rotation ensures 24/7 coverage without burnout
□ Test runbooks regularly to verify they still work
□ Escalate promptly when incident exceeds scope


References ​