Change Management & Postmortem Practices Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Change Management & Postmortem Practices flashcards as text
A postmortem identifies that a deployment pipeline lacked a staging environment, causing the bug to reach production directly. What type of action item should be created?
Answer: Create and enforce a mandatory staging environment that mirrors production configuration, with CI/CD gates that block production deployment until staging tests pass
The root cause is a missing technical control (staging environment). The correct action is to build that control and make it mandatory in the pipeline — a technical change that prevents the class of failure, not a process or documentation change.
Which of the following BEST describes the 'contributing factors' section of a postmortem?
Answer: A comprehensive list of conditions that allowed the incident to occur or worsen, including technical, process, and organizational factors — distinct from the single root cause
Contributing factors are the broader set of conditions that enabled or amplified the incident — they may include technical debt, missing monitoring, process gaps, or organizational pressures — and differ from the specific triggering cause.
Why is it important to track postmortem action items in a formal system (e.g., bug tracker) rather than just in the postmortem document?
Answer: Formal tracking enables assignment, prioritization, and accountability — ensuring action items are completed rather than becoming documentation artifacts that are never acted upon
Without formal tracking with owners and deadlines, postmortem action items frequently go unimplemented. A bug tracker creates visibility, accountability, and enables regular review of completion status across the team and organization.
A team is preparing a postmortem for a 4-hour outage that affected 30% of users. Which elements are MANDATORY in a high-quality postmortem?
Answer: Incident summary, timeline of events, root cause analysis, contributing factors, user impact, action items with owners and due dates, and lessons learned
A standard postmortem template includes: summary, timeline, root cause, contributing factors, impact, action items (with owners and due dates), and lessons learned — these elements enable learning and prevention.
What is the 'normalization of deviance' phenomenon, and why is it a significant risk in SRE change management?
Answer: When small deviations from safe practices are tolerated repeatedly without immediate negative consequences, teams gradually accept those deviations as normal, increasing systemic risk over time
Normalization of deviance, identified by sociologist Diane Vaughan studying the Challenger disaster, describes how accumulated small rule deviations that don't immediately cause harm gradually become accepted as normal, creating hidden systemic risk.
What does 'mean time between failures' (MTBF) measure, and why is it potentially misleading as a primary reliability metric?
Answer: MTBF measures the average time between incidents; it is misleading because two services with the same MTBF but very different incident durations (MTTR) will have vastly different availability
MTBF alone doesn't capture availability. A service with MTBF of 30 days but MTTR of 1 day has very different availability than one with MTBF of 30 days but MTTR of 1 minute. Availability = MTBF / (MTBF + MTTR).