← All SRE Flashcard Decks

On-Call Escalation and Runbook Design Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 On-Call Escalation and Runbook Design flashcards as text
  1. What is 'on-call rotation design,' and which team size and rotation frequency minimizes on-call fatigue?

    Answer: A minimum of 8 engineers in the rotation with a weekly rotation provides adequate coverage while ensuring each engineer is on-call less than once every two months, though actual minimum depends on service criticality and alert volume

    Google's SRE practice recommends at least 8 people in the rotation so that each person is on-call roughly every 8 weeks, providing adequate recovery time between rotations. Smaller teams need to either hire or reduce service scope.

  2. What is 'blameless escalation culture' and how does it impact escalation behavior?

    Answer: A culture where escalating an incident is treated as responsible and professional, not as an admission of incompetence — encouraging early escalation before incidents worsen rather than late escalation driven by fear of judgment

    When engineers fear being judged for escalating, they delay escalation — extending incidents, increasing user impact, and creating exactly the outcome that blame was supposedly trying to prevent. Blameless escalation makes early escalation the valued behavior.

  3. What is the 'load shedding' technique in on-call runbooks, and when should it be applied?

    Answer: Load shedding deliberately rejects a portion of incoming requests when the service is under extreme load to prevent total failure and maintain service quality for the requests that are accepted

    Load shedding is a controlled partial degradation strategy: reject lower-priority requests to protect the service's ability to serve higher-priority or core requests during overload, preventing a cascade to total failure.

  4. What does 'operational readiness review' (ORR) mean for a new on-call rotation member, and what should be verified?

    Answer: An ORR verifies that the new team member has required access, understands the services they will support, has read key runbooks, and has completed a shadow rotation before taking primary on-call responsibility

    An ORR for a new on-call member verifies operational readiness: access provisioned, runbooks read and understood, shadow rotation completed, and critical escalation paths known — ensuring they can effectively respond before holding primary responsibility.

  5. What is 'alert deduplication' in monitoring systems, and why is it important for on-call effectiveness?

    Answer: Alert deduplication groups multiple alerts with the same root cause into a single notification, preventing the on-call engineer from being flooded with hundreds of related alerts from a single incident

    Without deduplication, a single incident can generate hundreds of alerts (one from each affected service or host), burying the on-call engineer in notifications and making it hard to understand the scope and nature of the incident.

  6. What is the 'two-minute rule' in runbook design and what problem does it address?

    Answer: If a runbook step takes more than two minutes to read and understand, it should be broken into smaller, more focused steps — ensuring the responder can act quickly without reading paragraphs of context mid-incident

    Long, dense runbook steps force the responder to stop and read extensively during a high-pressure incident. Breaking instructions into small, scannable steps that can be understood and executed in under two minutes maintains momentum and reduces cognitive load.