Monitoring, Alerting & Troubleshooting Flashcards
7 cards from real PCA practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Monitoring, Alerting & Troubleshooting flashcards as text
An Alertmanager route has `group_wait: 30s` and `group_interval: 5m`. A second alert fires in the same group 10 seconds after the first. When will the second alert be sent?
Answer: Together with the first alert after the 30s group_wait expires
Alerts in the same group are batched together; the second alert joins the existing group and is sent when the group_wait of 30s expires.
Which Prometheus metric type is most appropriate for tracking the total number of HTTP requests served since process start?
Answer: Counter
Counters are monotonically increasing values ideal for cumulative totals like request counts.
A Prometheus alert has been in the `PENDING` state for longer than the configured `for` duration but has not transitioned to `FIRING`. What is the most likely cause?
Answer: The alert condition became false before the for-duration elapsed
If the alert condition resolves before the for-duration completes, the alert resets to inactive without ever firing.
What does the `absent()` function return when the time series passed to it exists and has a value?
Answer: An empty vector
absent() returns an empty vector when the time series is present, and a vector with value 1 only when the series is absent.
In Prometheus, what is the effect of setting `honor_labels: true` on a scrape configuration?
Answer: Conflicting labels from the target override Prometheus-assigned labels
With honor_labels: true, if the target exposes a label that conflicts with a Prometheus-set label, the target's label wins.
You notice Prometheus is dropping samples with the error 'out of order'. What is the most likely root cause?
Answer: Samples are arriving with timestamps older than the current TSDB head block
Prometheus TSDB requires samples to arrive in order; samples with timestamps behind the latest ingested timestamp for a series are rejected as out of order.
Which Alertmanager feature allows sending a notification only once for a group of related alerts, rather than once per alert?
Answer: Grouping
Grouping batches related alerts (e.g., same cluster or job) into a single notification to reduce alert fatigue.