← All AWS Flashcard Decks

Design Resilient Architectures 10 Flashcards

6 cards from real AWS practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Design Resilient Architectures 10 flashcards as text
  1. An Aurora Global Database has a primary cluster in us-east-1 and a secondary in eu-west-1. During a regional disaster drill, the team promotes the secondary to primary. After the primary region recovers, the team wants to restore the original topology. Which sequence correctly restores the original primary without data loss?

    Answer: Detach the recovered us-east-1 cluster, add it back as a secondary to the promoted eu-west-1 primary, allow replication to catch up, then perform a managed planned failover back to us-east-1.

    After an unplanned failover (promotion), the original primary is detached and becomes a standalone cluster. You must re-add it as a secondary (read replica) to the new global primary, let replication fully synchronize, and then use a managed planned failover to restore the desired topology — all without data loss. Restoring from snapshot introduces an RPO gap, Aurora does not support dual-primary, and DMS is not the correct tool for this topology switch.

  2. An SQS queue has a visibility timeout of 30 seconds and a message retention period of 4 days. A consumer crashes after receiving a message but before deleting it. The consumer is offline for exactly 32 seconds, then restarts. What is the state of the message?

    Answer: The message becomes visible again and is reprocessed by the consumer after 30 seconds.

    Once the visibility timeout (30 seconds) expires without a DeleteMessage call, the message automatically becomes visible in the queue again and is eligible for redelivery to any consumer — including the restarted one. SQS does not move messages to a DLQ simply for visibility timeout expiry; DLQ routing only occurs after the maxReceiveCount threshold is breached. The message is never deleted without an explicit DeleteMessage call, and retention governs maximum lifetime, not re-visibility behavior.

  3. A company uses Route 53 latency-based routing across three regions. They add a health check to each record. The us-west-2 endpoint fails its health check. A user in California (closest to us-west-2) makes a DNS query. What does Route 53 return?

    Answer: The record with the next-lowest latency for that user (e.g., us-east-1 or eu-west-1) that is currently healthy.

    When health checks are attached to latency-based routing records, Route 53 combines both policies: it evaluates latency rankings and then filters out unhealthy endpoints. The California user would normally get us-west-2, but since that record is unhealthy, Route 53 falls back to the next-best latency option that passes its health check. Route 53 never returns SERVFAIL for healthy alternates, and it does not send unhealthy IPs with TTL 0 — unhealthy records are simply excluded from the response set.

  4. A DynamoDB global table is configured across us-east-1 and ap-southeast-1. Both replicas accept writes (active-active). A network partition occurs: ap-southeast-1 can no longer reach us-east-1 for 8 minutes. Clients in Asia continue writing to the ap-southeast-1 replica. When connectivity is restored, how does DynamoDB resolve conflicting writes to the same item that occurred on both replicas during the partition?

    Answer: DynamoDB uses a last-writer-wins strategy based on the timestamp embedded in each write, with the write carrying the most recent timestamp winning.

    DynamoDB global tables use a last-writer-wins reconciliation model based on a server-side timestamp. When connectivity is restored after a partition, conflicting writes to the same item are resolved by keeping the write with the most recent timestamp — the application is not involved in conflict resolution. DynamoDB does not halt writes during partitions (that would sacrifice availability), does not treat either region as authoritative, and does not raise ConditionalCheckFailedExceptions for replication conflicts.

  5. An ECS service uses a FARGATE_SPOT capacity provider with a base of 0 and weight of 1, and a FARGATE capacity provider with a base of 2 and weight of 1. The service is running 10 tasks. AWS sends a two-minute Spot interruption notice for 4 running Fargate Spot tasks simultaneously. What is the MOST LIKELY outcome?

    Answer: ECS attempts to launch 4 replacement tasks using Fargate Spot or standard Fargate per capacity provider weights, but the service may temporarily drop below 10 tasks if capacity is insufficient.

    Fargate Spot interruptions give a two-minute notice, and ECS will attempt to start replacement tasks — but it does not guarantee they will be running before the interrupted tasks are reclaimed. Replacement tasks are placed according to capacity provider weights and available capacity; if Spot capacity is constrained (a common cause of interruptions), replacement tasks may land on standard Fargate but could take time to reach RUNNING state. ECS does not guarantee zero-drop replacements within two minutes, does not preemptively drain all tasks, and does not unconditionally move the whole service to standard Fargate.

  6. A Kinesis Data Stream has 10 shards. A consumer application using the Enhanced Fan-Out (EFO) API loses connectivity for 7 minutes and then reconnects. Each EFO shard iterator is valid for 5 minutes. What does the consumer receive when it calls SubscribeToShard after reconnecting?

    Answer: The consumer must call RegisterStreamConsumer or re-invoke SubscribeToShard with a fresh iterator position; it will receive data starting from the earliest available position it specifies (e.g., AFTER_SEQUENCE_NUMBER), potentially missing records only if they aged out of the 24-hour (or configured) retention window.

    Enhanced Fan-Out shard iterators expire after 5 minutes of inactivity. After a 7-minute disconnection, the iterator is expired and the consumer must call SubscribeToShard again with a new starting position (e.g., AFTER_SEQUENCE_NUMBER using the last successfully processed sequence number). Records are retained for 24 hours by default (up to 7 days with extended retention), so records produced during the outage are retrievable — none are lost unless they exceed the retention window. EFO does not maintain persistent consumer state for automatic resume, and there is no 'rebalancing' step triggered by an expired iterator.