Alerting Strategy and Noise Reduction Flashcards
7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Alerting Strategy and Noise Reduction flashcards as text
According to the Google SRE book, which question should every alert be able to answer 'yes' to in order to justify its existence?
Answer: Is this alert actionable—does a human need to do something in response right now?
The Google SRE book's core alerting principle states that every page-worthy alert must require immediate human action; alerts that don't need a human response right now should be tickets or removed entirely.
What is the 'Four Golden Signals' framework for monitoring and alerting in distributed systems?
Answer: Latency, Traffic, Errors, and Saturation—the key indicators of service health
The Four Golden Signals (Latency, Traffic, Errors, Saturation) defined in the Google SRE book provide a minimal but comprehensive framework for alerting on the health of any user-facing service.
How does error budget burn rate directly influence alerting strategy?
Answer: Burn rate determines the appropriate alert window: fast burn triggers short-window alerts; slow burn triggers long-window alerts
Fast error budget burn (e.g., 14x over 1 hour) warrants short-window urgent alerts, while slow burn (e.g., 1x over days) warrants longer-window lower-priority alerts—matching alert urgency to actual impact speed.
What is a 'phantom alert' in SRE monitoring practice?
Answer: An alert that fires based on stale or delayed metrics and resolves itself before an engineer can respond
Phantom alerts fire due to data pipeline delays or metric staleness, creating false urgency that resolves on its own—they erode trust in the alerting system and contribute to engineers ignoring alerts.
What is the recommended strategy for handling alert noise during a known, expected maintenance window?
Answer: Silence or suppress alerts for the affected services for the duration of the maintenance window
Silencing alerts for known maintenance windows prevents noise that would otherwise desensitize engineers, while ensuring the silence is time-bounded and explicitly tied to the maintenance event for accountability.
In a microservices architecture, what alerting anti-pattern does 'alert on every dependency failure' represent?
Answer: Cascading alerts, where a single upstream failure generates dozens of downstream service alerts
Alerting on every dependency failure causes cascading alerts where one upstream outage spawns dozens of downstream alerts, overwhelming on-call engineers with noise instead of pointing to the single root cause.
What is the benefit of using 'alert ownership' metadata (such as team labels) in an alerting system?
Answer: It ensures alerts are routed to the correct on-call team with the knowledge to resolve the issue
Alert ownership metadata ensures pages are routed to the team responsible for that service, reducing response time and preventing confusion about who should be investigating a given alert.