Mixed Deck — All SRE Topics Flashcards
100 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 20 Mixed Deck — All SRE Topics flashcards as text
In a service mesh, what is 'traffic mirroring' (shadow traffic), and how is it used to improve reliability?
Answer: Traffic mirroring sends a copy of live production requests to a shadow service version, allowing testing of new versions with real traffic patterns without affecting production users
Traffic mirroring (also called shadow deployments or dark launches) sends an asynchronous copy of production traffic to a shadow instance, allowing performance testing and behavior validation under real load without any risk to production users.
What is a 'rollback' in release engineering?
Answer: Reverting production to a previously known-good version
A rollback restores the system to a previous stable version when a new release causes problems.
What is 'throughput' vs. 'latency' in performance testing, and what is their typical trade-off relationship?
Answer: Throughput is the number of requests processed per second; latency is the time to process a single request. At low load, both can be optimized simultaneously, but at high load, higher throughput often comes with increased latency as queuing occurs
At low utilization, adding more requests doesn't significantly increase individual request latency. As utilization approaches capacity, queueing theory (Little's Law) predicts latency increases sharply — this is the hockey-stick latency curve observed in most systems.
What is 'synthetic monitoring' and how does it complement real user monitoring (RUM)?
Answer: Synthetic monitoring runs scripted probe transactions against the production service continuously; combined with RUM (which captures real user experience), it provides complete observability: synthetics detect issues when real user traffic is absent, RUM captures the actual diversity of user experiences
Synthetic monitoring provides 24/7 baseline measurements and immediate alerts even at zero user traffic. RUM captures the true diversity of user environments, devices, and geographies that synthetics cannot replicate — together they cover different monitoring blind spots.
What is 'horizontal vs. vertical scaling,' and when is each approach MOST appropriate?
Answer: Horizontal scaling adds more instances of a service; vertical scaling increases the resources (CPU, memory) of existing instances. Horizontal scaling is preferred for stateless services and enables near-infinite capacity; vertical scaling is simpler but has hardware limits and causes downtime during upgrades
Horizontal scaling (scale out) adds instances and distributes load — works well for stateless services but requires load balancing and may introduce distributed system complexity. Vertical scaling (scale up) is simpler but hits hardware ceilings and typically requires a restart.
A team defines a latency SLI as 'the proportion of homepage requests served in under 200 ms.' Which SLO statement is BEST aligned with this SLI?
Answer: 99% of homepage requests will be served in under 200 ms over a rolling 28-day window
An SLO must reference the same SLI metric and unit. The SLI is the proportion of requests under 200 ms, so the SLO should set a target percentage for that same proportion over a defined time window.
What is 'data sovereignty' in the context of disaster recovery, and how does it constrain DR architecture for regulated industries?
Answer: Data sovereignty laws require that certain data remain within specific geographic jurisdictions; DR backups and replicas must be stored in compliant locations, potentially preventing the use of the geographically distant DR sites that would otherwise provide the best resilience
Laws like GDPR (EU), data localization laws (Russia, China, India), and industry regulations (healthcare, finance) restrict where certain data can be stored — this can force DR to use in-country or in-region backup sites rather than the most geographically distant, resilient locations.
When a SRE professional encounters an unfamiliar challenge in distributed systems design, what is the recommended first course of action?
Answer: Research applicable standards, consult with subject matter experts, and document the approach
Professional practice requires a methodical approach to unfamiliar challenges: research the applicable standards, consult experts when needed, and document the reasoning for the chosen approach.
In PagerDuty or OpsGenie, what is an 'override' typically used for?
Answer: Temporarily assigning on-call duty to a different engineer for a defined time window
An override temporarily substitutes one engineer for another in the schedule, used for vacations, sick days, or planned absences without changing the permanent rotation.
What is 'mutual TLS' (mTLS) in a service mesh, and why is it important for microservices security?
Answer: mTLS requires both the client and server to present certificates, ensuring that only authenticated services can communicate — preventing impersonation attacks within the cluster
mTLS provides bidirectional authentication: both sides of every service-to-service connection verify the other's identity via certificate, preventing a compromised pod from impersonating a legitimate service or intercepting traffic.
Which log severity level should be used for events that indicate the system is in an unrecoverable state and will shut down?
Answer: CRITICAL/FATAL
CRITICAL or FATAL severity indicates an unrecoverable condition that forces the application to abort or shut down.
An on-call engineer receives a P1 alert at 2 AM but cannot determine the cause after 20 minutes of investigation. What is the CORRECT action?
Answer: Escalate to the secondary on-call or the designated expert as defined in the escalation policy, without waiting longer — unresolved P1 incidents require additional resources
P1 incidents have defined escalation timeouts in the incident management policy — waiting beyond them is a policy violation that risks extended user impact. Escalation is not failure; it is the defined process.
Which element should every runbook include to help engineers confirm a remediation worked?
Answer: Expected system behavior or metrics after the fix is applied
Verification steps tell the engineer what to check after a remediation to confirm the issue is resolved, preventing premature incident closure.
In the context of capacity planning, what is 'demand forecasting'?
Answer: Predicting future resource needs based on business and traffic growth trends
Demand forecasting uses historical trends, business growth projections, and seasonality to predict future infrastructure resource requirements.
What is 'alert grouping' in notification systems like PagerDuty or Alertmanager?
Answer: Combining multiple related alert firings into a single notification to reduce noise
Alert grouping aggregates multiple related alerts (e.g., the same error firing across 50 pods) into a single notification, preventing notification floods while preserving the information that an incident is occurring.
Which of the following best describes 'RED' method metrics in SRE observability?
Answer: Rate, Errors, Duration
The RED method focuses on Rate (requests per second), Errors (failed requests), and Duration (distribution of request latencies) for services.
Which of the following best describes the 'five whys' technique in a postmortem?
Answer: Iteratively asking 'why' about each cause until a systemic root cause is identified, typically requiring around five iterations
The five whys is a root cause analysis technique: start with the failure, ask why it happened, then ask why that cause happened, repeating until a fundamental systemic cause is uncovered — typically taking about five iterations.
What does 'policy as code' mean in infrastructure management?
Answer: Defining and enforcing governance rules using code rather than manual processes
Policy as code means encoding compliance, security, and operational rules as machine-readable code that can be automatically enforced and version-controlled.
What is the PRIMARY risk of over-automating processes that still require human judgment?
Answer: Automation can mask system anomalies by silently 'fixing' symptoms without addressing root causes
Blind automation of remediation can hide underlying problems by repeatedly correcting symptoms while the root cause continues to worsen.
What is 'business impact analysis' (BIA) and how does it drive DR investment decisions?
Answer: BIA quantifies the financial and operational impact of each service being unavailable, providing the data needed to justify DR investment by showing the cost of downtime versus the cost of recovery capabilities
BIA maps each service to its business impact when unavailable (revenue loss, regulatory penalties, customer attrition, reputational damage) and its recovery cost, enabling data-driven decisions about RTO/RPO targets and DR investment levels.