Reliability Principles & Service Level Objectives Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Reliability Principles & Service Level Objectives flashcards as text
Why does Google's SRE model argue that having a separate SRE team, rather than embedding reliability work in development teams, is beneficial?
Answer: A dedicated SRE team creates structural incentives to prioritize reliability — developers are motivated to ship features, while SREs are specifically accountable for stability and operational excellence
Google's SRE model creates a team whose primary incentive and success metric is reliability, counterbalancing development teams whose primary incentive is feature velocity. The structural separation aligns incentives with the organization's reliability goals.
An SLA promises customers 99.5% monthly availability with a 10% service credit for each additional 0.5% of downtime. The service experienced 5 hours of downtime in a 30-day month. Was the SLA breached?
Answer: Yes — 99.5% of 43,200 minutes allows only 216 minutes (3.6 hours) of downtime; 5 hours (300 minutes) exceeds this
30 days × 24 hours × 60 minutes = 43,200 minutes. 0.5% × 43,200 = 216 minutes (3.6 hours) allowed. 5 hours = 300 minutes, which exceeds 216 minutes, so the SLA is breached.
What is the key difference between an SLO and an SLA?
Answer: SLOs are internal targets that guide engineering decisions; SLAs are external contractual commitments with financial or legal consequences for breach
SLOs are internal targets that teams use to drive reliability work and gate releases. SLAs are contracts with customers — if the SLA is breached, there are consequences (credits, contract termination). SLOs are typically set stricter than SLAs to provide a buffer.
A reliability review finds that a service achieves 99.97% availability — well above its 99.9% SLO. The SRE team proposes loosening the SLO to 99.95%. What is the BEST argument in favor of this change?
Answer: The gap between actual reliability and the SLO target suggests the service is over-engineered for its current needs; loosening the SLO frees error budget for feature velocity without meaningfully impacting users
If a service consistently and significantly exceeds its SLO, it may be over-engineered relative to user needs, consuming engineering resources that could be used for features. Adjusting the SLO to reflect the actual user-acceptable threshold releases that budget.
Which of the following best describes 'toil' in the SRE context?
Answer: Manual, repetitive, automatable operational work that scales linearly with service load and does not produce lasting improvement
Toil is the specific category of work that is manual, repetitive, tactical (no enduring improvement), reactive, and scales proportionally with service growth — the opposite of engineering work that reduces future burden.
Why should SLOs be set based on what users actually need rather than on what the system can currently achieve?
Answer: Setting SLOs based on current capability locks in the status quo and may over-invest in reliability that users don't value, while user-need-based SLOs create the right reliability incentives
SLOs calibrated to current capability lock in technical debt and over-engineering simultaneously. User-need-based SLOs create a clear target: meet the minimum reliability users require, invest the rest in innovation.