SRE Foundation Certification Exam β Questions and Answers
Question 1: What role does continuous improvement play in on-call practices & runbooks for SRE certified professionals?
- It drives ongoing enhancement of practices, processes, and outcomes through systematic evaluation (Correct answer)
- It applies only to new professionals in their first year
- It focuses exclusively on cost reduction
- It is optional and only necessary during certification renewal
Correct answer: It drives ongoing enhancement of practices, processes, and outcomes through systematic evaluation
Continuous improvement is fundamental to professional practice in on-call practices & runbooks, involving regular evaluation, feedback integration, and process enhancement to maintain high standards.
Question 2: What is the 'load shedding' technique in on-call runbooks, and when should it be applied?
- Load shedding is the process of distributing load evenly across all available server instances to prevent any single instance from becoming a hotspot
- Load shedding deliberately rejects a portion of incoming requests when the service is under extreme load to prevent total failure and maintain service quality for the requests that are accepted (Correct answer)
- Load shedding refers to the practice of moving non-urgent on-call tasks to the next business day to reduce after-hours workload
- Load shedding is a power management technique used in data centers during peak electricity demand to reduce cooling load
Correct answer: Load shedding deliberately rejects a portion of incoming requests when the service is under extreme load to prevent total failure and maintain service quality for the requests that are accepted
Load shedding is a controlled partial degradation strategy: reject lower-priority requests to protect the service's ability to serve higher-priority or core requests during overload, preventing a cascade to total failure.
Question 3: What does 'Mean Time to Acknowledge' (MTTA) measure in SRE on-call operations?
- The average time between repeated occurrences of the same alert over a rolling 30-day window
- The average time from incident declaration to when a postmortem review is completed
- The average time from when an alert fires to when an engineer acknowledges receipt of the page (Correct answer)
- The average time from alert creation to when the alert condition is fully resolved
Correct answer: The average time from when an alert fires to when an engineer acknowledges receipt of the page
MTTA measures the speed of initial human response to alerts, serving as a KPI for on-call responsiveness and alerting system effectivenessβhigh MTTA may indicate alert fatigue or poor notification routing.
Question 4: What is 'incident commander' (IC) role and what are the TOP THREE responsibilities during a major incident?
- Contact affected customers, refund SLA credits, and brief the board of directors on the incident impact
- Monitor all metrics dashboards, approve all production changes, and write the final root cause analysis
- Write the postmortem, manage the on-call rotation, and deploy the hotfix that resolves the incident
- Coordinate all responders and communications, make decisions on mitigation actions and escalations, and maintain situational awareness across the incident β NOT to personally diagnose or fix the technical issue (Correct answer)
Correct answer: Coordinate all responders and communications, make decisions on mitigation actions and escalations, and maintain situational awareness across the incident β NOT to personally diagnose or fix the technical issue
The IC's value is coordination and decision-making authority, not technical execution. By delegating technical tasks, the IC maintains the bird's-eye view needed to make tactical decisions and keep the incident moving toward resolution.
Question 5: Which technique allows a team to test system behavior under 2Γ production load without generating real user traffic?
- Increasing CDN TTLs
- Feature flag rollout to 50% of users
- Shadow traffic (dark launch) mirroring (Correct answer)
- Enabling database read replicas
Correct answer: Shadow traffic (dark launch) mirroring
Shadow traffic mirrors real production requests to a parallel environment, allowing load testing at scale with real patterns without impacting users.
Question 6: What is the purpose of a 'blameless post-mortem' in on-call culture?
- To analyze what went wrong and improve systems without attributing personal fault (Correct answer)
- To calculate the financial cost of the incident
- To identify which engineer caused the incident so they can be disciplined
- To file a formal complaint against a vendor
Correct answer: To analyze what went wrong and improve systems without attributing personal fault
Blameless post-mortems focus on systemic causes and process improvements rather than individual blame, fostering a culture where engineers report problems honestly.
Question 7: A microservice shows memory usage growing 5% per week. Which action should the SRE prioritize?
- Restart the service weekly to reset memory
- Switch to a language with automatic memory management
- Immediately double memory on all instances
- Investigate for a memory leak while monitoring the growth trend (Correct answer)
Correct answer: Investigate for a memory leak while monitoring the growth trend
Steady growth suggests a memory leak; the correct response is to investigate the root cause while tracking the trend to determine urgency.
Question 8: What is 'cost-aware autoscaling,' and how does it balance reliability with cost efficiency?
- Cost-aware autoscaling integrates cost signals (e.g., spot instance availability and pricing) with performance signals to make scaling decisions that optimize for both reliability and cost β such as scaling out on cheap spot instances during low-priority batch work while reserving on-demand capacity for SLO-critical services (Correct answer)
- Cost-aware autoscaling is the default mode for all cloud autoscalers and requires no special configuration
- Cost-aware autoscaling eliminates autoscaling entirely in favor of fixed provisioning based on worst-case traffic estimates
- Cost-aware autoscaling reduces the maximum replica count to stay within a monthly budget, potentially sacrificing SLO compliance during traffic peaks
Correct answer: Cost-aware autoscaling integrates cost signals (e.g., spot instance availability and pricing) with performance signals to make scaling decisions that optimize for both reliability and cost β such as scaling out on cheap spot instances during low-priority batch work while reserving on-demand capacity for SLO-critical services
Cost-aware autoscaling differentiates between workloads β user-facing services get reliable on-demand capacity to protect SLOs, while batch/background workloads use spot instances when available for maximum cost efficiency.
Question 9: Which practice best reduces cognitive load for an on-call engineer during an incident?
- Sending all logs to a single stream
- Pre-writing runbooks with clear, numbered steps for known failure modes (Correct answer)
- Disabling non-critical alerts permanently
- Granting the on-call engineer full production access
Correct answer: Pre-writing runbooks with clear, numbered steps for known failure modes
Pre-written runbooks allow engineers to follow a structured response plan rather than improvising under stress, significantly reducing cognitive load.
Question 10: What does 'fault injection' mean in the context of resilience testing?
- Deploying broken code to staging environments
- Deliberately introducing errors, latency, or resource exhaustion into a running system (Correct answer)
- Blocking all external traffic to simulate a DDoS event
- Scanning source code for potential runtime exceptions
Correct answer: Deliberately introducing errors, latency, or resource exhaustion into a running system
Fault injection intentionally introduces controlled failures (network drops, CPU spikes, disk full) to observe how the system responds.
Question 11: What is the primary purpose of a log aggregation pipeline in an SRE context?
- To delete logs older than 30 days automatically
- To centralize logs from multiple sources for unified search and analysis (Correct answer)
- To compress logs to save disk space
- To convert structured logs into unstructured text
Correct answer: To centralize logs from multiple sources for unified search and analysis
Log aggregation centralizes logs from many services into a single platform, enabling unified querying and correlation.
Question 12: What is 'infrastructure as code' (IaC) and why is it critical for disaster recovery?
- IaC defines infrastructure in version-controlled code (Terraform, CloudFormation, Pulumi), enabling rapid, consistent recreation of the entire environment in a DR region from scratch, eliminating 'snowflake' servers that cannot be reproduced (Correct answer)
- IaC is only useful for initial provisioning; it has no role in disaster recovery scenarios where speed is critical
- IaC is a documentation standard that describes infrastructure in human-readable format for compliance audits
- IaC is a monitoring approach that automatically codes alerts when infrastructure anomalies are detected
Correct answer: IaC defines infrastructure in version-controlled code (Terraform, CloudFormation, Pulumi), enabling rapid, consistent recreation of the entire environment in a DR region from scratch, eliminating 'snowflake' servers that cannot be reproduced
IaC enables reproducible environment creation β in a disaster, running terraform apply or deploying a CloudFormation stack recreates the entire production environment in a new region in minutes, rather than hours of manual provisioning that may contain configuration errors.
Question 13: What is the primary purpose of setting Service Level Objectives (SLOs)?
- To replace Service Level Agreements (SLAs).
- To ensure 100% uptime.
- To align business and reliability expectations (Correct answer)
- To eliminate the need for monitoring.
Correct answer: To align business and reliability expectations
Service Level Objectives (SLOs) are targets for a service's reliability, defining an acceptable level of performance or availability. Their primary purpose is to create a shared understanding and alignment between the service provider (SRE team) and business stakeholders regarding what constitutes 'good enough' reliability. This alignment helps in making informed decisions about resource allocation, feature development, and operational priorities.
Question 14: Why should SLOs be set based on what users actually need rather than on what the system can currently achieve?
- Setting SLOs based on current capability locks in the status quo and may over-invest in reliability that users don't value, while user-need-based SLOs create the right reliability incentives (Correct answer)
- User-need-based SLOs are required by industry compliance frameworks like ISO 27001
- User surveys are the only valid input for SLO setting because engineers cannot estimate reliability
- System-capability-based SLOs are always lower than user-need SLOs, leading to SLA breaches
Correct answer: Setting SLOs based on current capability locks in the status quo and may over-invest in reliability that users don't value, while user-need-based SLOs create the right reliability incentives
SLOs calibrated to current capability lock in technical debt and over-engineering simultaneously. User-need-based SLOs create a clear target: meet the minimum reliability users require, invest the rest in innovation.
Question 15: An SRE team proposes that the service should have a separate SLO for mobile app users and desktop web users. What is the STRONGEST argument for this approach?
- Regulatory requirements mandate separate SLO tracking for different device types
- Mobile and desktop users may have different reliability expectations, latency tolerances, and failure modes β separate SLOs allow each to be optimized independently (Correct answer)
- Separate SLOs are simpler to calculate because each SLI has fewer dimensions
- Mobile users generate less revenue and therefore should have a lower reliability target
Correct answer: Mobile and desktop users may have different reliability expectations, latency tolerances, and failure modes β separate SLOs allow each to be optimized independently
Different user segments may have legitimately different reliability needs β mobile users may be more tolerant of latency but less tolerant of errors, or vice versa. Separate SLOs allow the team to optimize for each segment's actual needs.
Question 16: Which of the following best describes 'RED' method metrics in SRE observability?
- Response, Error, Dependency
- Reliability, Efficiency, Durability
- Rate, Errors, Duration (Correct answer)
- Requests, Events, Delays
Correct answer: Rate, Errors, Duration
The RED method focuses on Rate (requests per second), Errors (failed requests), and Duration (distribution of request latencies) for services.
Question 17: A change management policy requires a 48-hour review window for all production changes. An SRE argues this policy creates more risk than it reduces. What is the BEST argument supporting the SRE's position?
- Long review windows delay security patches and bug fixes, creating a larger window of vulnerability, and may incentivize engineers to batch changes in ways that increase blast radius (Correct answer)
- 48-hour reviews are only necessary for changes to customer-facing services, not internal infrastructure
- Review processes should be automated and do not need human review windows
- Engineers will simply bypass the review process if it is too slow, making it ineffective
Correct answer: Long review windows delay security patches and bug fixes, creating a larger window of vulnerability, and may incentivize engineers to batch changes in ways that increase blast radius
Fixed delay review windows create perverse incentives: engineers batch many changes together to minimize review overhead, creating larger, riskier deployments. They also delay critical security and stability fixes that need to ship immediately.
Question 18: What role does continuous improvement play in capacity planning & scaling for SRE certified professionals?
- It is optional and only necessary during certification renewal
- It focuses exclusively on cost reduction
- It applies only to new professionals in their first year
- It drives ongoing enhancement of practices, processes, and outcomes through systematic evaluation (Correct answer)
Correct answer: It drives ongoing enhancement of practices, processes, and outcomes through systematic evaluation
Continuous improvement is fundamental to professional practice in capacity planning & scaling, involving regular evaluation, feedback integration, and process enhancement to maintain high standards.
Question 19: What distinguishes 'overhead' from 'toil' in the SRE framework?
- Overhead has lasting value but no direct production impact, while toil is repetitive manual work scaling with service growth (Correct answer)
- Toil includes strategic planning; overhead is only operational work
- There is no meaningful distinction between the two terms
- Overhead always requires automation; toil is acceptable at any level
Correct answer: Overhead has lasting value but no direct production impact, while toil is repetitive manual work scaling with service growth
Overhead refers to administrative work (meetings, HR tasks) that has organizational value but doesn't scale with service load, unlike toil which grows proportionally.
Question 20: Which command initializes a Terraform working directory and downloads required providers?
- terraform setup
- terraform init (Correct answer)
- terraform get
- terraform install
Correct answer: terraform init
`terraform init` initializes the working directory, downloads provider plugins, and sets up the backend for state storage.
Question 21: An SRE is capacity planning for a batch processing system that runs nightly. What is the most cost-effective cloud strategy?
- Use spot/preemptible instances that scale up at night and terminate after completion (Correct answer)
- Keep a large fixed cluster running 24/7
- Use on-demand instances permanently sized for peak batch load
- Offload processing to a third-party SaaS tool
Correct answer: Use spot/preemptible instances that scale up at night and terminate after completion
Spot or preemptible instances at scale for short-duration batch jobs significantly reduce cost since the workload is fault-tolerant and time-flexible.
Question 22: What does 'mean time between failures' (MTBF) measure, and why is it potentially misleading as a primary reliability metric?
- MTBF measures the average time between incidents; it is misleading because two services with the same MTBF but very different incident durations (MTTR) will have vastly different availability (Correct answer)
- MTBF measures the time to deploy a fix after a failure is detected; it is misleading because it doesn't capture detection time
- MTBF measures the predicted hardware lifespan; it is misleading for software services because software doesn't wear out
- MTBF measures the frequency of deployment changes; it is misleading because not all changes cause failures
Correct answer: MTBF measures the average time between incidents; it is misleading because two services with the same MTBF but very different incident durations (MTTR) will have vastly different availability
MTBF alone doesn't capture availability. A service with MTBF of 30 days but MTTR of 1 day has very different availability than one with MTBF of 30 days but MTTR of 1 minute. Availability = MTBF / (MTBF + MTTR).
Question 23: Which practice ensures that a failed pipeline step stops downstream stages from running?
- Fast-fail (fail-fast) configuration where any failing step aborts the pipeline (Correct answer)
- Using feature flags to bypass failing steps
- Running tests only in the final stage
- Parallel execution of all stages
Correct answer: Fast-fail (fail-fast) configuration where any failing step aborts the pipeline
Fail-fast stops pipeline execution immediately on failure, preventing wasted compute and surfacing problems early.
Question 24: What does 'pipeline as code' mean in modern CI/CD practices?
- Using a GUI to configure pipeline steps
- Storing pipeline execution logs in the code repository
- Defining pipeline configuration in version-controlled files alongside application code (Correct answer)
- Writing pipeline logic in the application's main codebase
Correct answer: Defining pipeline configuration in version-controlled files alongside application code
Pipeline as code means storing pipeline definitions (e.g., Jenkinsfile, .gitlab-ci.yml) in version control for auditability and reproducibility.
Question 25: What is 'blameless escalation culture' and how does it impact escalation behavior?
- A culture where escalated incidents are never assigned a responsible party, ensuring no engineer faces consequences for their actions
- A culture where escalation decisions are made collectively by the team rather than by the on-call engineer individually
- A culture where all escalations require manager approval to prevent unnecessary interruptions
- A culture where escalating an incident is treated as responsible and professional, not as an admission of incompetence β encouraging early escalation before incidents worsen rather than late escalation driven by fear of judgment (Correct answer)
Correct answer: A culture where escalating an incident is treated as responsible and professional, not as an admission of incompetence β encouraging early escalation before incidents worsen rather than late escalation driven by fear of judgment
When engineers fear being judged for escalating, they delay escalation β extending incidents, increasing user impact, and creating exactly the outcome that blame was supposedly trying to prevent. Blameless escalation makes early escalation the valued behavior.
Question 26: Which of the following best describes a key competency required for chaos engineering & resilience in SRE practice?
- The ability to work independently without any oversight
- Strong analytical skills combined with effective communication and ethical judgment (Correct answer)
- Reliance on a single methodology for all situations
- Memorization of all relevant regulations without understanding context
Correct answer: Strong analytical skills combined with effective communication and ethical judgment
SRE professionals working in chaos engineering & resilience need analytical skills to assess situations, communication skills to convey findings, and ethical judgment to make sound decisions.
Question 27: What is the key difference between 'high availability' (HA) and 'fault tolerance' in system design?
- HA is achieved through geographic distribution; fault tolerance requires on-premises hardware only
- HA focuses on data durability; fault tolerance focuses on request speed
- HA minimizes downtime through redundancy and failover; fault tolerance allows a system to continue operating correctly even during active failures (Correct answer)
- HA applies only to databases; fault tolerance applies only to compute layers
Correct answer: HA minimizes downtime through redundancy and failover; fault tolerance allows a system to continue operating correctly even during active failures
HA aims to reduce outage duration via fast failover, while fault tolerance means the system continues operating (possibly at reduced capacity) without any perceptible disruption even while a component is failing.
Question 28: What is 'service versioning' in microservices, and what strategy allows multiple API versions to coexist without breaking existing consumers?
- Requiring all consumers to update to the latest API version within a 30-day deprecation window
- Running multiple API versions simultaneously (v1, v2) with routing by URL path or header, allowing consumers to migrate at their own pace while new features are available in v2 (Correct answer)
- Incrementing a version number in the service's deployment manifest to track which version is running in production
- Using semantic versioning (major.minor.patch) in the container image tag to distinguish between deployments
Correct answer: Running multiple API versions simultaneously (v1, v2) with routing by URL path or header, allowing consumers to migrate at their own pace while new features are available in v2
Running multiple API versions simultaneously with path-based routing (/v1/, /v2/) or header-based routing allows existing consumers to continue using the stable v1 API while new consumers adopt v2, enabling safe migration without coordinated cutover.
Question 29: What is the 'reliability hierarchy' in SRE, and what does it imply about incident priorities?
- Monitoring β Incident Response β Postmortem β Automation β Capacity Planning β each layer must be solid before the next adds value; an incident cannot be well-managed if monitoring is blind (Correct answer)
- SLA β SLO β SLI β Error Budget β each layer constrains the one below it in importance
- Hardware β Network β OS β Application β User β incidents should always be debugged from the bottom layer up
- Customer satisfaction β Features β Performance β Availability β reliability priorities should follow business value
Correct answer: Monitoring β Incident Response β Postmortem β Automation β Capacity Planning β each layer must be solid before the next adds value; an incident cannot be well-managed if monitoring is blind
The SRE reliability pyramid establishes that without solid monitoring, incident response is blind; without good incident response, postmortems lack data; without postmortems, automation lacks direction β each level depends on the foundation below it.
Question 30: Which principle emphasizes balancing innovation with reliability?
- 100% uptime policy
- Resource capping
- Load balancing
- Error budgeting (Correct answer)
Correct answer: Error budgeting
Error budgeting is a core SRE principle that explicitly balances the tension between innovation (deploying new features) and reliability (maintaining system stability). By defining an acceptable level of unreliability (the error budget), teams can use the remaining budget for experimentation and new deployments. If the budget is exhausted, focus shifts back to reliability work, ensuring a sustainable pace of development and balancing these two critical aspects.
Question 31: A service team wants to define SLIs for a batch data pipeline that processes files every night. Which SLI is MOST appropriate?
- Proportion of nightly batch runs that complete successfully within the defined processing window (e.g., 4 hours) (Correct answer)
- Time since the last successful batch run as reported by the scheduler
- Number of files processed per hour during peak load
- Average CPU utilization of the batch processing servers during runs
Correct answer: Proportion of nightly batch runs that complete successfully within the defined processing window (e.g., 4 hours)
For a batch pipeline, the key user concern is whether the batch completes on time and successfully. The proportion of runs completing within the processing window directly measures this user-visible concern.
Question 32: In a follow-the-sun on-call model, what is the main advantage?
- On-call responsibility transfers between regional teams to align with business hours (Correct answer)
- Incidents are only addressed during peak traffic hours
- All engineers are on-call simultaneously
- Engineers in one region handle all global incidents
Correct answer: On-call responsibility transfers between regional teams to align with business hours
Follow-the-sun rotations hand off on-call duty across time zones so engineers respond during their normal working hours, reducing night-time pages.
Question 33: What does 'policy as code' mean in infrastructure management?
- Using code reviews as the only approval gate for infrastructure changes
- Writing infrastructure cost budgets in spreadsheet formulas
- Defining and enforcing governance rules using code rather than manual processes (Correct answer)
- Automating HR policy distribution via configuration management
Correct answer: Defining and enforcing governance rules using code rather than manual processes
Policy as code means encoding compliance, security, and operational rules as machine-readable code that can be automatically enforced and version-controlled.
Question 34: What is 'right-sizing' in cloud capacity management?
- Scaling all services to the same instance type for consistency
- Matching instance or resource size to actual workload requirements to minimize waste (Correct answer)
- Provisioning at the maximum possible size for safety
- Choosing the cheapest available instance type
Correct answer: Matching instance or resource size to actual workload requirements to minimize waste
Right-sizing analyzes actual resource consumption and selects the instance type or size that meets performance requirements without over-provisioning.
Question 35: What is a 'rollback' in release engineering?
- Reverting production to a previously known-good version (Correct answer)
- Deploying a hotfix on top of the current version
- Running a load test before a new release
- Deleting failed deployment artifacts
Correct answer: Reverting production to a previously known-good version
A rollback restores the system to a previous stable version when a new release causes problems.
Question 36: An alert fires at 3 AM and the on-call engineer determines no user impact exists. What is the appropriate action?
- Immediately page the entire team to investigate
- Acknowledge, document the lack of impact, and create a ticket to tune the alert (Correct answer)
- Escalate to the VP of Engineering
- Silence the alert permanently without investigation
Correct answer: Acknowledge, document the lack of impact, and create a ticket to tune the alert
Non-impacting alerts should be acknowledged and documented, then a follow-up ticket created to reduce noise β silencing without investigation risks missing real issues.
Question 37: What is the 'risk of excessive caution' that SRE texts warn about?
- Over-cautious monitoring systems generate too many alerts, overwhelming the on-call team
- Excessive caution in deployment procedures causes developers to bypass safety checks entirely
- If a team is too conservative about reliability, they may under-invest in features and innovation, ultimately harming the business and the user experience more than occasional outages would (Correct answer)
- Being too careful about capacity planning leads to over-provisioning and unnecessary cloud costs
Correct answer: If a team is too conservative about reliability, they may under-invest in features and innovation, ultimately harming the business and the user experience more than occasional outages would
Extreme risk aversion in reliability engineering can starve a product of the innovation needed to remain competitive. An unreleased perfect service helps no users; an occasionally imperfect but continuously improving service may serve users better.
Question 38: Which metric type is most appropriate for tracking the total number of HTTP requests received since service start?
- Counter (Correct answer)
- Histogram
- Gauge
- Summary
Correct answer: Counter
A counter is a monotonically increasing value ideal for cumulative counts like total requests, errors, or bytes.
Question 39: In the context of SRE certification, what is the most important consideration when implementing release engineering & ci/cd?
- Completing implementation as quickly as possible regardless of quality
- Ensuring alignment with established standards, stakeholder needs, and best practices (Correct answer)
- Delegating all responsibilities to junior staff
- Minimizing documentation to save time
Correct answer: Ensuring alignment with established standards, stakeholder needs, and best practices
When implementing release engineering & ci/cd, SRE professionals must ensure alignment with industry standards and stakeholder needs. Hasty implementation without proper planning often leads to compliance issues and suboptimal outcomes.
Question 40: Which blast radius control technique limits a chaos experiment to only 10% of production pods?
- Blue-green traffic splitting
- Feature flag gating
- Canary deployment strategy
- Percentage-based targeting with a selector (Correct answer)
Correct answer: Percentage-based targeting with a selector
Percentage-based pod selectors (e.g., in LitmusChaos or Chaos Monkey) restrict the experiment scope to a fraction of the fleet.
SRE Foundation Certification Exam
The SRE Foundation Certification Exam validates an individual's understanding of the core principles and practices of Site Reliability Engineering (SRE).
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong β answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds