← All SRE Flashcard Decks

Disaster Recovery and Business Continuity Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Disaster Recovery and Business Continuity flashcards as text
  1. What is a 'GameDay' exercise in SRE, and what are its primary objectives?

    Answer: A GameDay is a planned disaster simulation where the team intentionally induces failure scenarios in a production or production-like environment to test incident response, runbooks, and system resilience

    GameDays deliberately introduce real failures to validate that the systems, runbooks, monitoring, and people all work together correctly under realistic conditions — finding gaps before real disasters do.

  2. In a multi-region cloud deployment, what is 'DNS failover,' and what are its limitations as a DR mechanism?

    Answer: DNS failover routes traffic to a secondary region by changing DNS records when the primary region's health check fails; limitations include DNS TTL propagation delay (minutes to hours) and client-side DNS caching that may prevent immediate failover

    DNS failover is widely used but has inherent delays — TTL values mean resolvers continue sending traffic to the failed primary for up to the TTL duration after the DNS record change, which can range from seconds (low TTL) to hours (high TTL) depending on the configuration.

  3. What is the 'shared fate' problem in DR architecture, and how do you design to avoid it?

    Answer: Shared fate occurs when the primary and DR system share a common failure domain (same data center, availability zone, cloud account, or dependency) that can take down both simultaneously; avoid by truly isolating DR infrastructure

    If the primary and DR systems share any critical component — the same power feed, same cloud region, same VPC, same DNS provider, or same network path — a failure of that shared component takes down both simultaneously, defeating the DR purpose.

  4. What is 'chaos engineering for DR validation,' and how does it differ from traditional DR testing?

    Answer: Chaos engineering for DR introduces failures in production continuously and unpredictably (like Netflix Chaos Monkey) rather than in scheduled exercises, building continuous confidence that DR mechanisms work under realistic conditions

    Traditional DR testing is periodic and scheduled — teams know it's coming and prepare. Chaos engineering runs continuously and without advance notice (at least to the systems), exposing real-world resilience gaps that scheduled tests often miss because teams anticipate and pre-fix them.

  5. A company's e-commerce platform has an RPO of 1 hour and RTO of 30 minutes for the product catalog service. Which backup and recovery architecture BEST meets these requirements?

    Answer: Synchronous replication to a warm standby in a second region with automated health-check-triggered failover, plus hourly snapshots to cross-region storage for point-in-time recovery

    A warm standby with continuous/near-continuous replication meets the 30-minute RTO (standby is pre-warmed, failover is automated). Hourly snapshots meet the 1-hour RPO requirement and provide point-in-time recovery for logical failures.

  6. What is 'data sovereignty' in the context of disaster recovery, and how does it constrain DR architecture for regulated industries?

    Answer: Data sovereignty laws require that certain data remain within specific geographic jurisdictions; DR backups and replicas must be stored in compliant locations, potentially preventing the use of the geographically distant DR sites that would otherwise provide the best resilience

    Laws like GDPR (EU), data localization laws (Russia, China, India), and industry regulations (healthcare, finance) restrict where certain data can be stored — this can force DR to use in-country or in-region backup sites rather than the most geographically distant, resilient locations.