โ† All SRE Flashcard Decks

Chaos Engineering & Resilience Flashcards

7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Chaos Engineering & Resilience flashcards as text
  1. What is the 'Swiss cheese model' as it applies to SRE resilience design?

    Answer: A layered defense model where multiple imperfect safeguards combine so no single failure causes a full outage

    The Swiss cheese model illustrates that layered defenses each have 'holes' (weaknesses), but when stacked, the holes rarely align, preventing complete system failure.

  2. What distinguishes a 'chaos engineering experiment' from random destructive testing?

    Answer: Chaos experiments have defined hypotheses, controlled scope, and measurable outcomes; random tests do not

    Legitimate chaos engineering is disciplined and scientific โ€” it starts with a hypothesis, limits blast radius, and measures specific outcomes to learn from each experiment.

  3. What is 'dependency mapping' and why is it a prerequisite for effective chaos experiments?

    Answer: Identifying all upstream and downstream service relationships so experiments target the right failure points

    Without knowing service dependencies, you cannot predict the blast radius of an experiment or design meaningful hypotheses about failure propagation.

  4. Which type of chaos experiment specifically validates that a service's circuit breaker trips correctly under sustained error conditions?

    Answer: Dependency failure simulation

    Simulating a dependency returning errors or timing out verifies that the calling service's circuit breaker opens correctly and stops cascading the failures downstream.

  5. What does 'graceful degradation' mean in resilience engineering?

    Answer: Shutting down non-critical features during high load so core functionality remains available

    Graceful degradation ensures that when parts of a system fail, the most critical features keep working even if secondary features become unavailable.

  6. In chaos engineering maturity models, what characterizes a 'Level 3' or advanced chaos practice?

    Answer: Fully automated continuous chaos integrated into CI/CD with real-time abort controls and self-healing validation

    Mature chaos programs run continuously and automatically as part of delivery pipelines, with automated safeguards and validation of self-healing mechanisms.

  7. What is the key difference between 'high availability' (HA) and 'fault tolerance' in system design?

    Answer: HA minimizes downtime through redundancy and failover; fault tolerance allows a system to continue operating correctly even during active failures

    HA aims to reduce outage duration via fast failover, while fault tolerance means the system continues operating (possibly at reduced capacity) without any perceptible disruption even while a component is failing.

Chaos Engineering & Resilience Flashcards โ€” SRE Study Cards with Answers