Incident and Event Response Flashcards
7 cards from real AWS practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Incident and Event Response flashcards as text
An EC2 instance is under memory pressure during an incident. Which CloudWatch metric would reveal this, and what is required to collect it?
Answer: MemoryUtilization — requires the CloudWatch Agent
MemoryUtilization is not published by default because AWS hypervisors cannot see guest OS memory; the CloudWatch Agent must be installed on the instance to collect and publish this metric.
Which AWS service uses machine learning to proactively detect anomalies in your application and infrastructure behavior without requiring you to configure manual thresholds?
Answer: Amazon DevOps Guru
Amazon DevOps Guru applies ML to your CloudWatch metrics, logs, and events to surface anomalous behavior and actionable recommendations without predefined alarm thresholds.
During an incident, your runbook must execute shell commands simultaneously on 50 EC2 instances. Which Systems Manager feature is designed for this use case?
Answer: Systems Manager Run Command
Systems Manager Run Command executes commands or scripts across a fleet of managed instances simultaneously without requiring SSH or RDP access.
Your incident response plan must automatically page on-call engineers via PagerDuty when a critical CloudWatch alarm fires. Which AWS service provides native integration for this?
Answer: AWS Systems Manager Incident Manager
Systems Manager Incident Manager supports native contact and escalation plan integrations with third-party tools like PagerDuty and OpsGenie through response plans.
What is the correct sequence of phases in the AWS incident management lifecycle?
Answer: Prepare → Detect → Respond → Recover → Learn
The AWS incident lifecycle starts with Prepare (runbooks and plans), then Detect, Respond, Recover, and Learn (post-incident analysis) — you must prepare before you can respond effectively.
An AWS Config automatic remediation keeps failing. Which CloudWatch metric should you examine to diagnose the failure count?
Answer: FailedRemediationAttempts
The FailedRemediationAttempts metric in AWS Config tracks how many remediation executions have failed, helping you identify and debug broken remediation configurations.
Which architecture detects AWS root account console logins and sends an SNS alert to the security team?
Answer: CloudTrail → CloudWatch Logs → metric filter → CloudWatch Alarm → SNS
The standard pattern is: create a CloudTrail trail delivering to CloudWatch Logs, create a metric filter matching ConsoleLogin events where userIdentity.type is Root, then alarm on that metric and notify via SNS.