Design Resilient Architectures 9 Flashcards
6 cards from real AWS practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Design Resilient Architectures 9 flashcards as text
An EC2 Auto Scaling group uses a termination lifecycle hook with a heartbeat timeout of 60 minutes and a default result of CONTINUE. The hook triggers a script that gracefully drains connections, but the script crashes silently after 5 minutes without ever sending a heartbeat or calling complete-lifecycle-action. What happens to the instance?
Answer: The instance stays in 'Terminating:Wait' for the full 60-minute timeout, then Auto Scaling applies the default result and terminates the instance.
Lifecycle hooks work on a timeout model, not a heartbeat-required model. The instance remains in the Terminating:Wait state until either complete-lifecycle-action is called (with CONTINUE or ABANDON) OR the heartbeat timeout expires. If the timeout expires with no action taken, Auto Scaling applies the DefaultResult — in this case CONTINUE — and proceeds with termination. Auto Scaling has no visibility into whether the hook script crashed; it only knows whether it received an explicit action call or the timeout elapsed.
An Aurora Global Database has its primary cluster in us-east-1 and a secondary cluster in eu-west-1. A catastrophic regional event destroys the us-east-1 primary — it is completely unreachable. The on-call team wants to promote eu-west-1 to primary. Which statement is most accurate about this recovery process?
Answer: Because the primary is destroyed, a managed planned failover cannot be initiated. The team must perform an unmanaged failover (detach and promote), which may incur data loss equal to the replication lag — typically under one second.
Aurora Global Database supports two failover paths: (1) Managed planned failover — requires the primary to be available so it can synchronize all pending WAL to the secondary before handing over. This achieves RPO=0 but cannot be used when the primary is destroyed. (2) Unmanaged failover — the team detaches the secondary from the global database and promotes it to a standalone cluster. This works even when the primary is gone, but any replication lag (typically <1 second but not guaranteed zero) represents potential data loss. Aurora does NOT automatically promote secondaries — manual intervention is always required for Aurora Global Database failover.
A DynamoDB Global Table has replicas in us-east-1 and ap-southeast-1. Due to a brief network partition, both regions accept a write to the same item (same partition key + sort key) within 50 milliseconds of each other — different attribute values, neither region aware of the other's write. Once the partition heals, what is the eventual state of the item in both regions?
Answer: DynamoDB uses a 'last writer wins' strategy based on wall-clock timestamps. Both regions converge to the version with the most recent timestamp; the other write is silently discarded.
DynamoDB Global Tables use a 'last writer wins' (LWW) reconciliation model based on item-level timestamps. When replication catches up after a network partition, each region evaluates the timestamps of conflicting writes. The item with the latest timestamp wins and is propagated to all replicas; the other write is silently discarded with no error surfaced to either writer. There is no attribute-level merge, no optimistic locking enforcement across regions (ConditionExpressions are evaluated locally), and no manual resolution step. Applications requiring true conflict-free semantics must implement application-level reconciliation or use conditional writes with caution.
A company runs a critical REST API on EC2 instances in private VPC subnets with no public IP addresses. They want to use Route 53 health checks to monitor individual backend instances and trigger DNS failover if an instance becomes unhealthy. A solutions architect proposes pointing Route 53 health checks directly at the private IP addresses. Why will this approach fail, and what is the correct alternative?
Answer: It will fail because Route 53 health checkers operate from AWS edge locations outside the VPC and cannot reach private IP addresses. The correct alternative is to create a Route 53 health check of type 'CloudWatch alarm' backed by a metric (such as EC2 StatusCheckFailed) from CloudWatch, which can observe private resources.
Route 53 health checkers are external AWS probers that initiate TCP/HTTP connections from outside the VPC. They have no network path to private IPs inside a VPC — security groups, NACLs, and the absence of internet routing all block them, and adding the health-checker IP range to a security group still won't help because there is no route to a private IP from outside the VPC. The correct pattern for private endpoints is to create a Route 53 health check of type 'CloudWatch alarm,' which delegates the health determination to a CloudWatch alarm. The alarm can be fed by metrics from CloudWatch Agent, ELB, or EC2 status checks — all of which can monitor private resources. This is the canonical AWS recommendation for VPC-internal health checks.
A Lambda function is triggered by an SQS queue using an event source mapping with batch size 10. The queue has a visibility timeout of 5 minutes and a dead-letter queue with maxReceiveCount=2. During a processing run, 9 messages in a batch succeed, but 1 message throws an unhandled exception. Lambda returns an error to the event source mapping. After two retries, the operations team finds all 10 messages in the DLQ, including the 9 that were successfully processed. What is the correct fix with the least operational overhead?
Answer: Enable partial batch response by setting FunctionResponseTypes to ['ReportBatchItemFailures'] on the event source mapping, and return a batchItemFailures list containing only the failed message ID.
By default, when a Lambda function returns an error, the SQS event source mapping treats the entire batch as failed and makes all messages visible again for retry. After maxReceiveCount attempts, all messages — including successfully processed ones — are routed to the DLQ. The purpose-built solution is ReportBatchItemFailures: the function explicitly returns a JSON body with a batchItemFailures array listing only the messageId values that failed. The event source mapping then deletes the successfully processed messages from the queue and leaves only the failed ones visible for retry. This avoids duplicate processing, prevents good messages from hitting the DLQ, and adds no per-message infrastructure overhead. Reducing batch size to 1 works but multiplies API call costs and reduces throughput significantly.
A company runs a stateful application across three Availability Zones behind an Application Load Balancer. Sessions are stored in-process on EC2 instances (not externalized). During a scale-in event, Auto Scaling begins terminating an instance that has 200 active sessions. The team sets the ALB deregistration delay to 300 seconds. A solutions architect reviews the setup and warns that active users will still lose their sessions. Why?
Answer: The deregistration delay keeps the instance registered for up to 300 seconds, allowing in-flight requests to complete, but it does not prevent Auto Scaling from terminating the EC2 instance before those 300 seconds elapse if no lifecycle hook is configured.
ALB deregistration delay (connection draining) stops the ALB from routing NEW requests to the deregistering target and gives existing in-flight requests up to 300 seconds to complete at the ALB layer. However, it has no effect on the EC2 Auto Scaling termination lifecycle. Without a termination lifecycle hook, Auto Scaling will issue the EC2 terminate command on its own schedule — which can be almost immediately after triggering scale-in — killing the instance and dropping all sessions regardless of what the ALB is waiting for. The correct architecture requires BOTH: (1) an ALB deregistration delay to drain the load balancer, AND (2) an Auto Scaling termination lifecycle hook that waits for draining to complete before allowing the instance to terminate. Ideally, sessions should also be externalized (ElastiCache, DynamoDB) so they survive instance replacement.