On-Call Escalation and Runbook Design Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 On-Call Escalation and Runbook Design flashcards as text
What is the 'follow-the-sun' on-call model, and what are its primary benefits and challenges?
Answer: Follow-the-sun distributes on-call coverage across geographically distributed teams in different time zones so that each team handles incidents during their business hours, reducing after-hours pages; the main challenge is handoff quality and context continuity
Follow-the-sun coverage ensures on-call engineers are never paged during their night by handing off responsibility as business hours shift around the globe — but requires excellent incident handoff processes to prevent information loss during transitions.
What is 'alert routing' in incident management, and why is it important to route alerts to the RIGHT team?
Answer: Alert routing directs each alert to the team with the knowledge and access to resolve it, reducing time-to-acknowledge and preventing incidents from sitting in an un-monitored queue while being routed to the wrong team
Routing alerts to the wrong team wastes critical response time — the team that receives the alert may spend time just figuring out who owns the service, delaying diagnosis and remediation.
A runbook step says: 'Check if the issue is in the database or application layer.' What makes this a POOR runbook instruction?
Answer: It is vague and non-actionable — a good runbook instruction would specify exactly which commands to run, which metrics to check, and how to interpret the results to make the determination
Vague instructions require domain knowledge that a stressed, sleep-deprived on-call engineer may not have at 3 AM. A good runbook specifies exact commands, expected outputs, and decision criteria so the procedure can be executed mechanically.
What is 'context-sensitive escalation' and how does it improve on simple time-based escalation?
Answer: Context-sensitive escalation considers the severity, service type, and current conditions (e.g., high-traffic event, recent deployment) to escalate differently depending on context — a high-traffic event may trigger faster escalation than a normal business day
Simple time-based escalation ('escalate if unresolved after 30 minutes') is inflexible. Context-sensitive escalation adjusts escalation urgency based on incident severity, business context (peak traffic periods), recent changes, and affected user tier.
What should a runbook include for an 'escalation decision tree' to be effective?
Answer: Clear decision points based on observable conditions (e.g., 'if mitigation not found in 20 min OR user impact >10%'), explicit escalation targets by name/role, contact information, and what information to share when escalating
An effective escalation decision tree removes ambiguity: the conditions for escalation are observable facts (not judgment calls), the escalation target is specific (not 'the database team'), and the information to communicate is specified (not assumed).
What is 'incident commander' (IC) role and what are the TOP THREE responsibilities during a major incident?
Answer: Coordinate all responders and communications, make decisions on mitigation actions and escalations, and maintain situational awareness across the incident — NOT to personally diagnose or fix the technical issue
The IC's value is coordination and decision-making authority, not technical execution. By delegating technical tasks, the IC maintains the bird's-eye view needed to make tactical decisions and keep the incident moving toward resolution.