← All SRE Flashcard Decks

Service Mesh and Microservices Reliability Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Service Mesh and Microservices Reliability flashcards as text
  1. What is a 'dead letter queue' (DLQ) in asynchronous microservices, and why is it important for reliability?

    Answer: A DLQ captures messages that fail to process after a maximum retry count, preserving them for manual inspection or reprocessing rather than losing them permanently

    DLQs prevent message loss by capturing persistently failing messages for later analysis and reprocessing, ensuring that processing failures do not cause silent data loss in event-driven systems.

  2. What is the 'saga pattern' in microservices, and what reliability problem does it solve?

    Answer: The saga pattern manages multi-step distributed transactions by defining compensating transactions for each step, ensuring that a failure midway through can be rolled back in a consistent manner without requiring distributed locks

    Sagas provide distributed transaction consistency in microservices without two-phase commit or distributed locks by defining compensating actions (rollbacks) that undo completed steps when a later step fails.

  3. How does 'request hedging' improve tail latency in microservices, and when should it be used carefully?

    Answer: Request hedging sends the same request to multiple backends simultaneously after a short delay and uses the first response, reducing p99 latency at the cost of additional backend load

    Hedged requests reduce tail latency by sending a second (or more) copy of the request to an alternative backend after a short delay (e.g., at p50 latency), using whichever response arrives first — trading increased load for reduced p99/p999 latency.

  4. What is 'service versioning' in microservices, and what strategy allows multiple API versions to coexist without breaking existing consumers?

    Answer: Running multiple API versions simultaneously (v1, v2) with routing by URL path or header, allowing consumers to migrate at their own pace while new features are available in v2

    Running multiple API versions simultaneously with path-based routing (/v1/, /v2/) or header-based routing allows existing consumers to continue using the stable v1 API while new consumers adopt v2, enabling safe migration without coordinated cutover.

  5. What is 'distributed tracing,' and which W3C standard protocol enables trace context propagation across service boundaries?

    Answer: Distributed tracing tracks a request as it flows through multiple services using a trace ID propagated in request headers; the W3C TraceContext standard defines the 'traceparent' and 'tracestate' headers for interoperability

    Distributed tracing reconstructs the full path of a request across services using trace IDs propagated in HTTP headers. The W3C TraceContext standard (traceparent/tracestate headers) provides vendor-neutral propagation for OpenTelemetry and other frameworks.

  6. A microservice team is designing a new API. Which approach BEST supports the reliability principle of graceful degradation?

    Answer: Return cached or default responses when downstream dependencies fail, rather than propagating errors to the caller — ensuring the API remains partially functional even when dependencies are unavailable

    Graceful degradation means the system continues to provide partial functionality when dependencies fail, serving cached, default, or reduced-feature responses rather than failing completely.