← All SRE Flashcard Decks

Disaster Recovery and Business Continuity Flashcards

6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Disaster Recovery and Business Continuity flashcards as text
  1. What is the difference between RTO (Recovery Time Objective) and RPO (Recovery Point Objective)?

    Answer: RTO is the maximum acceptable time to restore service after a disaster; RPO is the maximum acceptable amount of data loss measured in time (how old the most recent backup can be)

    RTO (Recovery Time Objective) defines how long the business can tolerate being down. RPO (Recovery Point Objective) defines how much data loss is acceptable — if RPO is 1 hour, backups must occur at least every hour.

  2. What is an 'active-active' disaster recovery architecture, and how does it differ from 'active-passive'?

    Answer: Active-active runs identical service instances in multiple regions simultaneously, all serving traffic; active-passive keeps a standby region that is not serving traffic until failover is triggered

    Active-active (multi-region live traffic) provides zero-downtime failover and distributes load, but requires conflict resolution for distributed writes. Active-passive (standby warm/cold) is simpler and cheaper but has failover time proportional to how 'warm' the standby is.

  3. Why is it critical to regularly TEST disaster recovery procedures rather than just documenting them?

    Answer: DR procedures drift from reality as systems change — untested plans frequently fail during actual disasters, and regular testing reveals gaps, validates RTOs, and ensures the team has practiced the steps

    Systems change continuously. DR procedures that weren't tested against the current production environment often fail — a backup that wasn't monitored may be corrupt, a failover script may reference decommissioned infrastructure, or the RTO assumption may have been unrealistic.

  4. A company's DR plan declares a 99.99% availability SLA. Their current backup strategy is daily backups to a single S3 bucket in the same AWS region. What is the MOST critical gap in this plan?

    Answer: Daily backups create an RPO of up to 24 hours, and a single-region backup does not protect against a regional AWS outage — both are incompatible with 99.99% availability

    99.99% availability ≈ 52 minutes of downtime per year. A 24-hour RPO and single-region backup is incompatible — a regional failure would cause data loss of up to 24 hours and recovery time far exceeding 52 minutes.

  5. What is a 'runbook' in the context of disaster recovery, and what makes it operationally effective?

    Answer: A runbook is a documented, step-by-step procedure for responding to specific failure scenarios; it is effective when it is concrete (exact commands), tested, maintained, and executable by someone unfamiliar with the system

    An effective runbook contains exact commands (not just concepts), has been tested in a real or realistic environment, is kept up to date with system changes, and can be followed by an on-call engineer who did not write it — typically at 3 AM under stress.

  6. What is the difference between a 'backup' and 'replication' as data protection strategies?

    Answer: Backups are point-in-time snapshots that protect against accidental deletion and data corruption; replication copies live data to another location in near real-time and protects against site failures but not data corruption

    Replication mirrors live data near-instantly (low RPO) but cannot protect against data corruption or accidental deletion since the corruption is replicated too. Backups capture a known-good state at a point in time, providing recovery from logical failures, but have higher RPO.