← All AWS Flashcard Decks

Design Resilient Architectures 7 Flashcards

6 cards from real AWS practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Design Resilient Architectures 7 flashcards as text
  1. A financial application uses DynamoDB Global Tables with regions in us-east-1 and eu-west-1. During a network partition lasting 90 seconds, both regions accept writes to the same item. After the partition heals, what is the conflict resolution outcome?

    Answer: DynamoDB uses last-writer-wins based on wall-clock time, with the most recently written item winning globally

    DynamoDB Global Tables use a last-writer-wins conflict resolution strategy based on wall-clock timestamps at the time of each write. The item version with the highest timestamp across all replicas wins and is propagated globally — no application-level intervention is required. There is no ConflictException or attribute merging; you cannot designate a custom version field for this purpose. This is why clock skew across regions can produce surprising outcomes.

  2. An SQS-based processing pipeline has a visibility timeout of 30 seconds. A consumer polls a message, begins processing, but the Lambda function cold start plus processing consistently takes 45 seconds. The dead-letter queue maxReceiveCount is set to 3. What is the most operationally correct fix that avoids premature DLQ routing?

    Answer: Call the ChangeMessageVisibility API mid-processing to extend the visibility timeout dynamically

    The correct solution is calling ChangeMessageVisibility during processing to extend the timeout before it expires. This prevents the message from becoming visible again (causing duplicate delivery and incrementing the receive count toward the DLQ threshold). Increasing maxReceiveCount only allows more retries — it doesn't prevent premature re-delivery. Increasing retention only delays expiry, not re-delivery. FIFO queues do not automatically extend visibility timeouts; they still respect the configured value.

  3. A multi-AZ RDS for PostgreSQL cluster has automated backups enabled with a 7-day retention period. A developer accidentally runs DROP TABLE on a critical table at 2:47 PM. The last automated snapshot completed at 3:00 AM. What is the most precise recovery method to minimize data loss?

    Answer: Use Point-in-Time Recovery (PITR) to restore the instance to 2:46 PM, then export and re-import only the affected table

    RDS Point-in-Time Recovery (PITR) uses automated backups combined with transaction logs to restore a DB instance to any second within the retention window — in this case to 2:46 PM, just before the DROP. After restoring to a new instance, you can export only the affected table and import it into the production instance, minimizing data loss to seconds. Restoring from the 3:00 AM snapshot loses nearly 12 hours of data. Read replicas replicate all DDL statements including DROP TABLE with minimal lag — they would also have lost the table. Enhanced Monitoring captures OS metrics, not query results or table data.

  4. You are designing a Route 53 active-active failover for an API across us-east-1 and us-west-2 using latency-based routing. A health check monitors each region's ALB. The us-east-1 ALB health check begins failing because of a misconfigured security group blocking Route 53 health checkers, even though the application itself is healthy. What happens to traffic?

    Answer: Route 53 marks us-east-1 unhealthy and routes all traffic to us-west-2, but continues re-evaluating every 30 seconds

    Route 53 marks the us-east-1 endpoint as unhealthy when its health check fails, and all traffic is shifted to us-west-2. Route 53 does continuously re-evaluate health checks (every 30 seconds for standard checks), so once the security group is corrected, traffic will resume routing to us-east-1. The 18% threshold is a real Route 53 concept for calculated health checks, not standard endpoint checks. Route 53 does not ignore health check failures for latency-based records — the health check is enforced regardless of routing policy type.

  5. An ECS Fargate service runs 10 tasks behind an ALB. You set the deployment configuration to minimumHealthyPercent=50 and maximumPercent=200. During a rolling deployment, the ALB target group deregistration delay is 300 seconds. The new task version starts returning 502 errors. What is the MOST LIKELY reason ECS does not automatically roll back?

    Answer: The deployment circuit breaker was not enabled on the ECS service, so ECS has no mechanism to detect and act on task failures

    ECS rolling deployments do not automatically roll back by default. The deployment circuit breaker must be explicitly enabled on the ECS service (with rollback=true) for ECS to detect consecutive task failures and automatically revert to the last stable task definition. Without the circuit breaker, ECS will continue attempting to replace tasks even if new ones are failing. CodeDeploy blue/green is one path to automated rollback but not the only one — the ECS-native circuit breaker is the direct answer here. The deregistration delay and minimumHealthyPercent are valid considerations but not the root cause of missing rollback.

  6. A company runs a stateful WebSocket application on EC2 instances behind a Network Load Balancer (NLB). During an AZ failure, clients with active WebSocket connections to instances in the failed AZ experience connection drops. The NLB has cross-zone load balancing DISABLED. After the AZ recovers, which statement about NLB behavior is accurate?

    Answer: The NLB automatically re-registers instances in the recovered AZ and resumes routing new connections to them without any configuration change

    When an AZ recovers, an NLB automatically resumes routing new connections to healthy, registered instances in that AZ — no manual re-registration or configuration change is required. Instances that were registered before the failure remain registered; the NLB's health checks re-evaluate them after recovery. Cross-zone load balancing being disabled means traffic from each AZ's NLB node only routes to targets in that AZ, but does not affect automatic recovery of a restored AZ. The 60-second DNS flapping protection described in option C is not an NLB behavior. Option D is fabricated — toggling cross-zone load balancing is not a recovery mechanism.