โ† All SRE Flashcard Decks

Chaos Engineering & Resilience Flashcards

7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Chaos Engineering & Resilience flashcards as text
  1. What is 'network partition testing' designed to validate in a distributed system?

    Answer: That services handle split-brain scenarios and communication loss between nodes gracefully

    Network partition tests verify that distributed services handle scenarios where nodes cannot communicate, preventing data inconsistency or full outages.

  2. Which chaos engineering principle states that experiments should be 'scientific' with a defined hypothesis?

    Answer: Hypothesis-driven experimentation principle

    Chaos engineering follows the scientific method: form a hypothesis about system behavior, inject failure, and measure whether the hypothesis holds.

  3. What is a 'runbook' and why is it critical to resilience during chaos experiments?

    Answer: A documented set of procedures for responding to and recovering from specific failure scenarios

    Runbooks provide step-by-step recovery instructions so on-call engineers can respond consistently and quickly when chaos experiments reveal failure modes.

  4. In resilience architecture, what does 'bulkhead isolation' prevent?

    Answer: Cascading failures from one service component spreading to unrelated components

    Bulkhead isolation (borrowed from ship hull design) partitions resources so that a failure in one pool cannot exhaust resources needed by another.

  5. What does 'mean time to recovery' (MTTR) measure in the context of chaos engineering outcomes?

    Answer: The average time taken to restore normal service after a failure is detected

    MTTR measures how quickly a team can restore service after an incident, and chaos engineering aims to reduce it by improving detection and recovery procedures.

  6. Which tool is commonly used to inject chaos into Kubernetes workloads using CRDs (Custom Resource Definitions)?

    Answer: LitmusChaos

    LitmusChaos uses Kubernetes-native CRDs and operators to define, schedule, and execute chaos experiments on pods, nodes, and network paths.

  7. What is the difference between 'proactive' and 'reactive' resilience practices?

    Answer: Proactive practices identify and fix weaknesses before incidents occur; reactive practices respond after failures happen

    Proactive resilience (chaos experiments, DR drills) finds weaknesses before users are impacted; reactive resilience (incident response, postmortems) handles failures after they occur.