← All SRE Flashcard Decks

On-Call Practices & Runbooks Flashcards

7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 On-Call Practices & Runbooks flashcards as text
  1. What does 'alert fatigue' describe in an on-call context?

    Answer: Desensitization to alerts caused by excessive noisy or false-positive pages

    Alert fatigue occurs when engineers receive so many low-quality or false-positive alerts that they begin ignoring or dismissing pages, risking missed real incidents.

  2. Which element should every runbook include to help engineers confirm a remediation worked?

    Answer: Expected system behavior or metrics after the fix is applied

    Verification steps tell the engineer what to check after a remediation to confirm the issue is resolved, preventing premature incident closure.

  3. A service has been paging on-call engineers 15 times per shift on average. According to SRE principles, what should be done?

    Answer: Investigate and eliminate the sources of noise through alert tuning or automation

    High page volume indicates poor alert hygiene; SRE practice calls for eliminating noise through tuning thresholds, improving automation, and fixing root causes.

  4. What is the main risk of having a single 'super-engineer' who is always on-call for critical systems?

    Answer: Single-person on-call creates a bus factor risk and leads to unsustainable burnout

    Depending on one person creates a single point of failure — if that person is unavailable or leaves, response capability collapses, and the workload causes burnout.

  5. What distinguishes a 'playbook' from a 'runbook' in some SRE organizations?

    Answer: Playbooks provide high-level strategy and decision trees; runbooks provide specific step-by-step procedures

    Some teams use 'playbook' to mean a higher-level guide with decision logic, while 'runbook' refers to a specific ordered set of executable steps for a known failure mode.

  6. Which of the following best describes a 'severity level' in incident management?

    Answer: A classification that describes the impact of an incident on users and business operations

    Severity levels (e.g., SEV1–SEV4) categorize incidents by their scope of user impact and business consequence, guiding response urgency and escalation.

  7. Why is it important for runbooks to include links to relevant dashboards and logs?

    Answer: To save the on-call engineer time locating diagnostic data during a time-sensitive incident

    Direct links to dashboards and log queries eliminate the time an on-call engineer spends searching for the right views during an incident, accelerating diagnosis.