Observability & Logging Flashcards
7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Observability & Logging flashcards as text
Which of the following best describes 'RED' method metrics in SRE observability?
Answer: Rate, Errors, Duration
The RED method focuses on Rate (requests per second), Errors (failed requests), and Duration (distribution of request latencies) for services.
What distinguishes a 'push' model from a 'pull' model in metrics collection?
Answer: In push, services send metrics to the collector; in pull, the collector scrapes targets
In a push model, applications actively send metrics to a backend (e.g., StatsD, InfluxDB), while in a pull model, the backend scrapes endpoints (e.g., Prometheus).
A service emits logs with a correlation ID. What is the primary benefit of including this ID?
Answer: It allows logs from different services handling the same request to be linked together
A correlation ID ties together all log entries across services for a single request, enabling end-to-end request tracing through logs.
Which OpenTelemetry component is responsible for receiving, processing, and exporting telemetry data?
Answer: Collector
The OpenTelemetry Collector receives telemetry from SDKs, applies processing pipelines, and exports to backends like Jaeger, Prometheus, or Elastic.
In log management, what is a 'retention policy' and why is it important for SRE?
Answer: A configuration that defines how long logs are stored before deletion, balancing cost and compliance
Retention policies define storage duration for logs, balancing the cost of storage against compliance requirements and incident investigation needs.
What problem does 'log rotation' solve in host-based logging?
Answer: It prevents log files from growing indefinitely and consuming all disk space
Log rotation periodically archives or deletes old log files to prevent unbounded disk consumption on the host.
An SRE team wants to alert when p99 latency exceeds an SLO threshold. Which metric type should they instrument?
Answer: Histogram or Summary
Histograms and Summaries capture latency distributions, enabling percentile calculations like p99 needed for SLO-based alerting.