โ† All SRE Flashcard Decks

Chaos Engineering & Resilience Flashcards

7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Chaos Engineering & Resilience flashcards as text
  1. What is 'latency injection' used to test in a microservices architecture?

    Answer: How upstream callers handle slow dependencies, including timeout and retry behavior

    Latency injection slows downstream calls artificially to verify that clients implement correct timeouts, retries, and fallbacks rather than hanging indefinitely.

  2. What is the purpose of an 'abort condition' in a chaos experiment?

    Answer: To automatically halt the experiment if system health degrades beyond a safe threshold

    Abort conditions (sometimes called 'circuit breakers' for experiments) stop chaos injection automatically when key health metrics exceed acceptable bounds.

  3. In the context of SRE, what does 'resilience' fundamentally mean?

    Answer: A system's ability to withstand and recover from failures while continuing to serve users acceptably

    Resilience is not about preventing all failures but about ensuring the system degrades gracefully and recovers quickly so user impact is minimized.

  4. Which practice involves deliberately taking down entire availability zones in production to test failover?

    Answer: Region evacuation drills

    Region or AZ evacuation drills redirect all traffic away from one zone to verify that the remaining zones can absorb the load and failover works as designed.

  5. What is 'chaos as code' and what benefit does it provide?

    Answer: Writing chaos scenarios in YAML/JSON so experiments are version-controlled, repeatable, and reviewable

    Defining chaos experiments as code (e.g., in Git) enables peer review, change tracking, and consistent reproduction of experiments across environments.

  6. Which concept describes an architecture where each microservice has its own database to prevent shared data-layer failures?

    Answer: Database-per-service pattern

    The database-per-service pattern prevents one service's database issues from cascading into failures for other services that would share the same data store.

  7. During a chaos experiment, an engineer notices the system's error rate has spiked to 40%. What should they do first?

    Answer: Trigger the abort condition and restore the system to steady state immediately

    A 40% error rate far exceeds safe thresholds; the correct action is to halt the experiment and recover the system before further damage occurs.