Chaos Engineering & Resilience Flashcards
7 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Chaos Engineering & Resilience flashcards as text
What is 'network partition testing' designed to validate in a distributed system?
Answer: That services handle split-brain scenarios and communication loss between nodes gracefully
Network partition tests verify that distributed services handle scenarios where nodes cannot communicate, preventing data inconsistency or full outages.
Which chaos engineering principle states that experiments should be 'scientific' with a defined hypothesis?
Answer: Hypothesis-driven experimentation principle
Chaos engineering follows the scientific method: form a hypothesis about system behavior, inject failure, and measure whether the hypothesis holds.
What is a 'runbook' and why is it critical to resilience during chaos experiments?
Answer: A documented set of procedures for responding to and recovering from specific failure scenarios
Runbooks provide step-by-step recovery instructions so on-call engineers can respond consistently and quickly when chaos experiments reveal failure modes.
In resilience architecture, what does 'bulkhead isolation' prevent?
Answer: Cascading failures from one service component spreading to unrelated components
Bulkhead isolation (borrowed from ship hull design) partitions resources so that a failure in one pool cannot exhaust resources needed by another.
What does 'mean time to recovery' (MTTR) measure in the context of chaos engineering outcomes?
Answer: The average time taken to restore normal service after a failure is detected
MTTR measures how quickly a team can restore service after an incident, and chaos engineering aims to reduce it by improving detection and recovery procedures.
Which tool is commonly used to inject chaos into Kubernetes workloads using CRDs (Custom Resource Definitions)?
Answer: LitmusChaos
LitmusChaos uses Kubernetes-native CRDs and operators to define, schedule, and execute chaos experiments on pods, nodes, and network paths.
What is the difference between 'proactive' and 'reactive' resilience practices?
Answer: Proactive practices identify and fix weaknesses before incidents occur; reactive practices respond after failures happen
Proactive resilience (chaos experiments, DR drills) finds weaknesses before users are impacted; reactive resilience (incident response, postmortems) handles failures after they occur.