← All SRE Flashcard Decks

Error Budgets & SLI/SLO Management Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Error Budgets & SLI/SLO Management flashcards as text
  1. A product manager argues that the engineering team should target 100% availability to maximize customer satisfaction. How should an SRE respond?

    Answer: 100% availability is unachievable and an SLO near 100% leaves no error budget for safe deployments or experiments

    100% availability is theoretically impossible — all systems have some failure rate. An SLO approaching 100% also leaves no error budget, paralyzing engineering teams who need budget to safely release changes.

  2. What does a 'burn rate' of 6x mean in the context of SLO alerting?

    Answer: The service is consuming its error budget six times faster than the rate that would exhaust it exactly at the end of the SLO window

    Burn rate is the ratio of the current error consumption rate to the rate that would exactly exhaust the budget by the end of the window. A 6x burn rate means the budget will be gone in 1/6 of the remaining window.

  3. Which of the following is the BEST description of an SLI (Service Level Indicator)?

    Answer: A quantitative measure of a specific aspect of service behavior, expressed as a ratio or rate

    An SLI is a quantitative measurement — a ratio of good events to total events, a request latency percentile, or similar metric that reflects service quality from the user's perspective.

  4. When reviewing quarterly error budget consumption, an SRE notices that 80% of budget was consumed by a single botched deployment. What organizational change would BEST prevent this in the future?

    Answer: Implement progressive delivery (canary or blue/green deployments) with automated rollback triggered by SLI degradation

    Progressive delivery strategies limit blast radius by routing only a fraction of traffic to the new version, with automated rollback when SLIs degrade — directly addressing the root cause of large budget consumption from deployments.

  5. A team operates a service with a 99.9% availability SLO over a rolling 30-day window. In the last 30 days, the service experienced 2 hours of full outage and 3 hours of partial degradation where 50% of requests failed. How much error budget was consumed?

    Answer: 2 hours full outage + 1.5 hours equivalent degradation = 3.5 hours out of 43.8 minutes budget (far exceeded)

    Both full outages and partial degradations consume error budget proportionally. 50% failure rate for 3 hours equals 1.5 hours of full-outage-equivalent budget. Total: 3.5 hours against a budget of only ~43.8 minutes — massively over budget.

  6. Which of the following metrics would make the POOREST SLI for an e-commerce checkout service?

    Answer: CPU utilization of the checkout servers

    CPU utilization is a system-internal metric that does not directly reflect user experience. Users don't observe CPU load — they observe whether checkout succeeds and how fast it is.