← All SRE Flashcard Decks

Service Mesh and Microservices Reliability Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Service Mesh and Microservices Reliability flashcards as text
  1. What is 'eventual consistency' in an event-driven microservices system, and how does it affect API design?

    Answer: After a write, reads may temporarily return stale data until all replicas and downstream event consumers converge; APIs should clearly document this behavior and clients should handle inconsistency gracefully

    In event-driven systems, processing an event and updating all downstream views takes time. APIs must be designed with this in mind: clients should not assume immediately consistent reads after writes, and APIs should communicate their consistency guarantees.

  2. What is the 'strangler fig pattern' and how does it support safe migration from a monolith to microservices?

    Answer: The strangler fig pattern incrementally replaces monolith functionality by routing specific requests to new microservices while the monolith still handles the rest, gradually 'strangling' the monolith until it can be retired

    Named after the strangler fig tree that grows around a host tree and eventually replaces it, this pattern gradually redirects specific functionality to new microservices via a routing layer, allowing safe incremental migration with the monolith as a fallback.

  3. What is 'API gateway' rate limiting, and what are the two main rate limiting algorithms used in production?

    Answer: Rate limiting caps the number of requests a client can make in a time window; the two main algorithms are token bucket (allows short bursts up to a capacity) and leaky bucket (smooths request rate to a constant output rate)

    Token bucket allows bursting (clients can use accumulated tokens for short spikes) while leaky bucket enforces a smooth, constant output rate. Both are used in API gateways, with token bucket being more common for API rate limiting.

  4. What is 'contract testing' in a microservices context, and why does it provide better reliability assurance than integration testing alone?

    Answer: Contract testing verifies that a service's API matches the expectations of its actual consumers, catching breaking changes before deployment without requiring all services to run together in a test environment

    Consumer-driven contract testing (e.g., Pact framework) captures the exact API expectations of each consumer and verifies providers against those expectations independently, enabling fast, reliable detection of breaking changes without complex full-stack test environments.

  5. What is 'chaos engineering,' and how does it relate to the reliability of microservices in production?

    Answer: Chaos engineering proactively injects controlled failures into production (or staging) systems to discover weaknesses in resilience before they manifest as user-facing incidents

    Chaos engineering, pioneered by Netflix's Chaos Monkey, systematically introduces controlled failures (node kills, network latency, dependency failures) to validate that microservices systems are resilient and that failover mechanisms actually work under realistic conditions.

  6. What is 'API pagination' and why is it important for the reliability of microservices that return large data sets?

    Answer: API pagination returns data in bounded chunks (pages) rather than all at once, preventing out-of-memory errors, timeout failures, and downstream overwhelm when result sets grow unexpectedly large

    Without pagination, a query that returns millions of records can exhaust service memory, cause database query timeouts, and overwhelm downstream consumers — making pagination a fundamental reliability requirement for any API where result set size is unbounded.