Disaster Recovery and Business Continuity Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Disaster Recovery and Business Continuity flashcards as text
What is 'service degradation mode' in business continuity planning, and how does it differ from full service outage?
Answer: Service degradation mode provides reduced functionality (e.g., read-only mode, cached responses, limited features) when some components are unavailable, keeping the service partially operational rather than completely down
Degradation mode trades full functionality for continued availability — for example, an e-commerce site in read-only mode can still show products and accept wishlists even if the checkout system is down, which is better than a complete outage.
What is 'infrastructure as code' (IaC) and why is it critical for disaster recovery?
Answer: IaC defines infrastructure in version-controlled code (Terraform, CloudFormation, Pulumi), enabling rapid, consistent recreation of the entire environment in a DR region from scratch, eliminating 'snowflake' servers that cannot be reproduced
IaC enables reproducible environment creation — in a disaster, running terraform apply or deploying a CloudFormation stack recreates the entire production environment in a new region in minutes, rather than hours of manual provisioning that may contain configuration errors.
What is the 'recovery hierarchy' concept in business continuity, and how should SREs use it to prioritize restoration order?
Answer: The recovery hierarchy ranks services by their criticality to core business function and dependencies, ensuring foundational services (authentication, networking, databases) are restored before dependent services (application tiers, reporting)
Restoring services in dependency order prevents wasted effort — an application server is useless without its database; a dashboard is useless without the services it monitors. The hierarchy maps dependencies to a restoration sequence.
What is 'mean time to recover' (MTTR) in a DR context, and what are the most impactful ways to reduce it?
Answer: MTTR in DR is the time from disaster declaration to service restoration; it is reduced most by automating failover steps, maintaining runbooks, pre-warming standby infrastructure, and regularly testing and measuring recovery time
MTTR in DR is dominated by detection time, decision time, and execution time. Automation eliminates manual execution steps; pre-warming eliminates provisioning time; practiced runbooks eliminate confusion and decision-making time.
What is 'business impact analysis' (BIA) and how does it drive DR investment decisions?
Answer: BIA quantifies the financial and operational impact of each service being unavailable, providing the data needed to justify DR investment by showing the cost of downtime versus the cost of recovery capabilities
BIA maps each service to its business impact when unavailable (revenue loss, regulatory penalties, customer attrition, reputational damage) and its recovery cost, enabling data-driven decisions about RTO/RPO targets and DR investment levels.
What is the 'pilot light' DR strategy in cloud environments, and when is it the MOST appropriate choice?
Answer: The pilot light strategy keeps a minimal core of DR infrastructure running (database replication, core networking) while other components are not provisioned until needed, balancing low ongoing cost with faster recovery than cold standby
Pilot light keeps the most expensive and slowest-to-provision components running (databases, with replication) while faster-to-provision components (app servers, load balancers) are provisioned only during failover — optimizing for cost efficiency while maintaining reasonable RTO.