Alerting Strategy and Noise Reduction Flashcards
7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Alerting Strategy and Noise Reduction flashcards as text
What is 'alert fatigue' in the context of SRE?
Answer: A state where on-call engineers become desensitized to alerts due to excessive or low-quality notifications
Alert fatigue occurs when engineers receive so many alerts—especially false positives—that they begin ignoring or dismissing them, increasing the risk of missing real incidents.
Which alerting philosophy does Google SRE recommend as a best practice for page-worthy alerts?
Answer: Alert on symptoms that directly indicate user-visible impact
SRE best practice is to alert on symptoms (user-visible impact) rather than causes, since causes can be diagnosed after the alert fires and symptom-based alerts have lower false-positive rates.
What is a 'dead man's switch' (watchdog) alert in SRE monitoring?
Answer: An alert that fires when it STOPS receiving expected periodic signals from a system
A dead man's switch alert fires when it stops receiving a regular heartbeat signal, ensuring that alerting system failures or complete service outages are detected even when no explicit error is generated.
In SRE alerting, what does 'precision' of an alert measure?
Answer: The proportion of alert firings that correspond to actual incidents requiring action
Alert precision measures the fraction of alert firings that are true positives (real incidents), with low precision indicating many false positives that contribute to alert noise.
What is the primary distinction between a 'page' alert and a 'ticket' alert in SRE practice?
Answer: Page alerts require immediate human action; ticket alerts can be addressed during business hours
Page alerts indicate situations requiring immediate human response (urgent, action-required now), while ticket alerts represent issues that can wait for the next business day without significant user impact.
What is 'alert flapping' and why is it problematic?
Answer: When an alert repeatedly transitions between firing and resolved states in a short time window
Alert flapping occurs when a metric oscillates around a threshold, causing rapid on/off transitions that generate excessive notifications without indicating a stable incident requiring intervention.
Which of the following is a key indicator that an alert should be removed or tuned down in severity?
Answer: The alert consistently fires and resolves without any engineer taking action
If an alert regularly fires and auto-resolves without requiring engineer action, it is a false positive that adds noise without value and should be tuned or removed to reduce alert fatigue.