← All SRE Flashcard Decks

Reliability Principles & Service Level Objectives Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Reliability Principles & Service Level Objectives flashcards as text
  1. What is the 'reliability hierarchy' in SRE, and what does it imply about incident priorities?

    Answer: Monitoring → Incident Response → Postmortem → Automation → Capacity Planning — each layer must be solid before the next adds value; an incident cannot be well-managed if monitoring is blind

    The SRE reliability pyramid establishes that without solid monitoring, incident response is blind; without good incident response, postmortems lack data; without postmortems, automation lacks direction — each level depends on the foundation below it.

  2. A service team wants to define SLIs for a batch data pipeline that processes files every night. Which SLI is MOST appropriate?

    Answer: Proportion of nightly batch runs that complete successfully within the defined processing window (e.g., 4 hours)

    For a batch pipeline, the key user concern is whether the batch completes on time and successfully. The proportion of runs completing within the processing window directly measures this user-visible concern.

  3. What is the relationship between service reliability and feature velocity in the SRE model?

    Answer: They are balanced through the error budget: when the budget is healthy, teams can move fast; when it is exhausted, reliability work takes priority over features

    The error budget is the mechanism that balances these competing priorities dynamically. Healthy budget = release freely; exhausted budget = pause and fix. Neither always wins — the budget decides.

  4. Which property of a good SLO measurement window is MOST important for catching slow reliability degradations?

    Answer: Using a rolling window (e.g., trailing 30 days) rather than a fixed calendar window, so degradation is continuously tracked

    Rolling windows continuously reflect the current reliability trajectory. A fixed calendar window can hide a slow degradation that started late in one period and continues into the next — the calendar reset discards accumulated history.

  5. An SRE team proposes that the service should have a separate SLO for mobile app users and desktop web users. What is the STRONGEST argument for this approach?

    Answer: Mobile and desktop users may have different reliability expectations, latency tolerances, and failure modes — separate SLOs allow each to be optimized independently

    Different user segments may have legitimately different reliability needs — mobile users may be more tolerant of latency but less tolerant of errors, or vice versa. Separate SLOs allow the team to optimize for each segment's actual needs.

  6. What is the 'risk of excessive caution' that SRE texts warn about?

    Answer: If a team is too conservative about reliability, they may under-invest in features and innovation, ultimately harming the business and the user experience more than occasional outages would

    Extreme risk aversion in reliability engineering can starve a product of the innovation needed to remain competitive. An unreleased perfect service helps no users; an occasionally imperfect but continuously improving service may serve users better.