← All AWS Flashcard Decks

Design Resilient Architectures 8 Flashcards

6 cards from real AWS practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Design Resilient Architectures 8 flashcards as text
  1. A financial services company runs a multi-tier application across three AWS Regions. Their RPO is 15 minutes and RTO is 30 minutes. The primary Region experiences a full outage. Which DR strategy meets BOTH objectives at the LOWEST cost?

    Answer: Pilot light in a second Region with core databases replicated via Aurora Global Database, and pre-defined runbooks to scale compute within the RTO window

    Pilot light with Aurora Global Database satisfies both constraints: Aurora Global DB provides sub-second RPO (well within 15 min) and its managed failover completes in under 1 minute, leaving ample time within the 30-minute RTO to scale compute from a minimal pilot-light footprint. Warm standby at 25% capacity also works technically but costs more than pilot light. Active-active exceeds requirements and costs significantly more. Cold standby cannot reliably meet a 30-minute RTO for a multi-tier application requiring full stack rebuild.

  2. An e-commerce platform uses SQS standard queues to decouple order processing. During a flash sale, a processing Lambda function throws throttling errors and messages are reprocessed multiple times, causing duplicate orders. The team has already set the visibility timeout to 6x the Lambda timeout. What is the MOST operationally efficient fix?

    Answer: Implement idempotency in the Lambda function using a DynamoDB conditional write keyed on the SQS MessageId before processing each order

    The root cause is that Lambda throttling causes messages to become visible again and be reprocessed. The correct fix is idempotency at the application layer: a conditional DynamoDB write (e.g., PutItem with a ConditionExpression checking the MessageId doesn't already exist) ensures that even if a message is delivered multiple times, only the first delivery successfully writes the order. FIFO queues reduce duplicates but at 300 TPS (or 3,000 with batching) throughput limits, which would worsen the flash sale problem. Setting max receive count to 1 causes messages to go to the DLQ on first failure rather than retrying. Reducing concurrency limits worsens throughput without fixing the duplicate-processing problem.

  3. A company uses an ALB in front of an Auto Scaling group. During a deployment, they use a blue/green strategy where the new (green) target group receives 10% of traffic via weighted target groups. After 10 minutes, CloudWatch metrics show the green group has a 0.5% error rate vs. 0.1% for blue. The deployment pipeline automatically promotes to 100% green if no alarms fire. Which action BEST prevents a bad deployment from fully promoting while minimizing manual intervention?

    Answer: Create a CloudWatch Composite Alarm that combines a metric alarm on the green target group's 5XX error rate exceeding 0.3% AND a missing-data alarm, then wire it to an EventBridge rule that calls a Lambda to revert the ALB weights

    A CloudWatch Composite Alarm on the green target group's 5XX error rate directly monitors the right signal (green-only errors, not aggregate), and wiring it to an EventBridge → Lambda automated rollback provides fast, fully automated remediation without manual steps. Critically, combining it with a missing-data alarm prevents a monitoring gap from allowing silent failures through. CodeDeploy with Linear10PercentEvery1Minute doesn't apply here — the deployment is already in progress with weighted target groups on an ALB, not a CodeDeploy-managed fleet. The Athena approach introduces 5+ minutes of latency and architectural complexity. Route 53 health checks operate at the load-balancer level and cannot distinguish error rates between target groups behind the same ALB.

  4. A serverless application uses API Gateway → Lambda → DynamoDB. The Lambda function reads an item, modifies it in memory, and writes it back. Under high concurrency, the team observes data corruption due to lost updates. They want the STRONGEST consistency guarantee with MINIMAL code changes. Which approach is correct?

    Answer: Use DynamoDB's UpdateItem with a ConditionExpression checking the item's version attribute, and increment the version on each write (optimistic locking)

    The problem is a classic read-modify-write race condition. DynamoDB's UpdateItem with a ConditionExpression (optimistic locking using a version attribute) is the correct solution: if two Lambdas read the same version, only one write succeeds; the other receives a ConditionalCheckFailedException and must retry. This is atomic and requires only adding a version attribute and a condition to the existing UpdateItem call. DynamoDB Streams with last-write-wins introduces data loss by design. Global tables with strongly consistent reads solve cross-region staleness, not concurrent write conflicts within the same Region. SDK retries with backoff only reduce collision probability — they do not prevent lost updates under sufficiently high concurrency.

  5. A company's architecture uses Kinesis Data Streams with 10 shards, consumed by a Lambda function. During a shard iterator expiration event caused by a 12-hour processing backlog, all iterator positions are lost. After restoring, the team wants to prevent this and ensure no records older than 24 hours are skipped. Which combination of settings achieves this?

    Answer: Increase the Kinesis data retention period to at least 48 hours, enable Lambda's on-failure destination to SQS, configure bisect-on-error, and set iterator starting position to AT_TIMESTAMP for recovery

    The core problem is that a 12-hour backlog caused iterator expiration (default Kinesis retention is 24 hours, but the iterator itself expires after 5 minutes of inactivity if not renewed — a Lambda processing backlog can exhaust this). Increasing retention to 48+ hours ensures records from 24+ hours ago are still in the stream. Using AT_TIMESTAMP as the recovery iterator position lets the team resume from exactly 24 hours ago without reprocessing everything. Bisect-on-error isolates problematic batches, and an SQS on-failure destination captures poison-pill records without blocking the stream. TRIM_HORIZON in option A would reprocess all retained records from the beginning, likely re-creating the backlog. Enhanced Fan-Out (option B) solves throughput, not iterator expiration or recovery position. Firehose (option D) changes the architecture entirely and doesn't use Kinesis Data Streams.

  6. An organization uses AWS Organizations with SCPs. A central security team applies an SCP to an OU that denies 's3:DeleteBucket'. A developer in a member account has an IAM policy with 's3:*' and is the bucket owner. The developer attempts to delete a bucket and receives an Access Denied error. The security team removes the SCP. The deletion still fails. What is the MOST LIKELY reason?

    Answer: The bucket has S3 Object Lock enabled in Compliance mode with a retention period that has not yet expired

    After the SCP is removed, the developer's 's3:*' IAM policy should permit deletion — unless an S3-layer control prevents it. S3 Object Lock in Compliance mode is enforced by S3 independently of IAM and SCPs: even the bucket owner (and even root) cannot delete a bucket or overwrite/delete locked objects until the retention period expires. No IAM policy, no SCP removal, and no AWS Support request can bypass Compliance mode. IAM propagation delays (option A) are real but sub-second in practice, not blocking after an SCP change. A bucket policy with an explicit Deny (option C) is plausible but would have also blocked deletion while the SCP was in place — the question implies deletion only started failing after the SCP was applied, meaning the bucket policy is not the new variable. Option D is fabricated — there is no such AWS waiting period.