Performance Testing and Load Management Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Performance Testing and Load Management flashcards as text
What is 'database query optimization' in the context of SRE performance management, and what is the most impactful first step?
Answer: Analyzing slow query logs and using EXPLAIN ANALYZE to identify queries missing indexes or performing full table scans — adding targeted indexes often reduces query time from seconds to milliseconds
Missing indexes are the most common cause of slow queries — a full table scan of a 10-million-row table takes seconds; an indexed lookup takes milliseconds. EXPLAIN ANALYZE reveals whether indexes are being used and which operations are expensive.
What is 'cache hit rate' and why is it a critical performance metric for content delivery systems?
Answer: Cache hit rate is the percentage of requests served from the cache without going to the origin; a high hit rate (e.g., 95%+) dramatically reduces origin load and latency, while a low hit rate means most traffic bypasses the cache and hits the origin
A cache hit rate of 95% means 95% of requests are served instantly from cache at low cost; only 5% reach the origin. A hit rate of 50% means half of all traffic hits the origin, potentially doubling origin load compared to the cached baseline.
What is 'horizontal vs. vertical scaling,' and when is each approach MOST appropriate?
Answer: Horizontal scaling adds more instances of a service; vertical scaling increases the resources (CPU, memory) of existing instances. Horizontal scaling is preferred for stateless services and enables near-infinite capacity; vertical scaling is simpler but has hardware limits and causes downtime during upgrades
Horizontal scaling (scale out) adds instances and distributes load — works well for stateless services but requires load balancing and may introduce distributed system complexity. Vertical scaling (scale up) is simpler but hits hardware ceilings and typically requires a restart.
What is 'request tracing' in performance debugging, and how does it differ from logging?
Answer: Request tracing follows a specific request through all services in its execution path, capturing timing data for each operation; logging records events without linking them to specific requests across service boundaries
Distributed tracing correlates all operations across all services for a single request using a trace ID, showing the call graph and timing. Logs are per-service event records with no inherent cross-service correlation without additional instrumentation.
What is 'SLO-based alerting,' and how does it improve alert quality compared to threshold-based alerting?
Answer: SLO-based alerting fires based on the rate of SLO budget consumption rather than crossing absolute thresholds — alerting only when there is a meaningful risk to the error budget, reducing false positives from transient spikes
SLO-based alerting uses burn rate and budget consumption to trigger alerts only when there's a real risk of budget exhaustion — avoiding false alarms from brief spikes that don't meaningfully consume the monthly budget.
What is 'memory leak' detection in production services, and what metrics are MOST useful for identifying it?
Answer: A memory leak causes continuous memory growth over time; it is identified by monotonically increasing memory usage metrics that do not decrease after garbage collection or request completion, and often correlates with increasing GC pause times
The signature of a memory leak in production metrics is memory usage that grows continuously over time without plateauing, combined with increasing GC pressure (more frequent GC, longer GC pauses) as the JVM or runtime attempts to free memory that cannot be freed.