SRE Foundation Certification Exam β Questions and Answers
Question 1: What should an SLO be tied to?
- Business objectives and user impact (Correct answer)
- Infrastructure cost
- Codebase complexity
- Deployment frequency
Correct answer: Business objectives and user impact
Service Level Objectives (SLOs) should always be tied directly to what matters most to the business and its users. They define the acceptable performance and reliability from the perspective of the end-user experience and critical business functions. This ensures that SRE efforts are focused on improving aspects that directly contribute to user satisfaction and overall business success, making them impactful and relevant.
Question 2: What is the purpose of a 'blameless post-mortem' in on-call culture?
- To analyze what went wrong and improve systems without attributing personal fault (Correct answer)
- To identify which engineer caused the incident so they can be disciplined
- To calculate the financial cost of the incident
- To file a formal complaint against a vendor
Correct answer: To analyze what went wrong and improve systems without attributing personal fault
Blameless post-mortems focus on systemic causes and process improvements rather than individual blame, fostering a culture where engineers report problems honestly.
Question 3: What is 'toil' in SRE terminology as it relates to on-call work?
- Repetitive, manual, and automatable operational work that scales with service growth (Correct answer)
- Any work that improves system reliability
- Writing documentation and runbooks
- Performing capacity planning exercises
Correct answer: Repetitive, manual, and automatable operational work that scales with service growth
Toil is repetitive manual work tied to running a production service that provides no enduring value and grows with the system β a primary target for automation.
Question 4: A postmortem identifies that a deployment pipeline lacked a staging environment, causing the bug to reach production directly. What type of action item should be created?
- Require engineers to manually test all changes on their local development machines before merging
- Add a code review requirement that specifically reviews for the type of bug that caused the incident
- Document the incident in the team wiki to ensure future engineers are aware of the risk
- Create and enforce a mandatory staging environment that mirrors production configuration, with CI/CD gates that block production deployment until staging tests pass (Correct answer)
Correct answer: Create and enforce a mandatory staging environment that mirrors production configuration, with CI/CD gates that block production deployment until staging tests pass
The root cause is a missing technical control (staging environment). The correct action is to build that control and make it mandatory in the pipeline β a technical change that prevents the class of failure, not a process or documentation change.
Question 5: Which tool is useful for tracking incidents and changes?
- MS Paint.
- YouTube.
- Jira or ServiceNow for traceability (Correct answer)
- Email only.
Correct answer: Jira or ServiceNow for traceability
Tools like Jira and ServiceNow are invaluable for tracking incidents and changes due to their comprehensive capabilities. They provide centralized platforms for logging, categorizing, assigning, and monitoring the entire lifecycle of incidents, problems, and changes. This ensures clear traceability, accountability, and provides historical data essential for analysis and continuous improvement in IT operations.
Question 6: What role does continuous improvement play in release engineering & ci/cd for SRE certified professionals?
- It is optional and only necessary during certification renewal
- It drives ongoing enhancement of practices, processes, and outcomes through systematic evaluation (Correct answer)
- It focuses exclusively on cost reduction
- It applies only to new professionals in their first year
Correct answer: It drives ongoing enhancement of practices, processes, and outcomes through systematic evaluation
Continuous improvement is fundamental to professional practice in release engineering & ci/cd, involving regular evaluation, feedback integration, and process enhancement to maintain high standards.
Question 7: What is the primary purpose of setting Service Level Objectives (SLOs)?
- To replace Service Level Agreements (SLAs).
- To align business and reliability expectations (Correct answer)
- To ensure 100% uptime.
- To eliminate the need for monitoring.
Correct answer: To align business and reliability expectations
Service Level Objectives (SLOs) are targets for a service's reliability, defining an acceptable level of performance or availability. Their primary purpose is to create a shared understanding and alignment between the service provider (SRE team) and business stakeholders regarding what constitutes 'good enough' reliability. This alignment helps in making informed decisions about resource allocation, feature development, and operational priorities.
Question 8: What is the SRE benefit of using declarative IaC over imperative scripts?
- Declarative tools require no provider credentials
- Imperative scripts cannot be stored in version control
- Declarative code runs significantly faster than imperative scripts
- Declarative IaC describes desired state, making the system responsible for achieving and maintaining it (Correct answer)
Correct answer: Declarative IaC describes desired state, making the system responsible for achieving and maintaining it
Declarative IaC lets you specify what you want rather than how to get there, enabling tools to handle idempotency and state reconciliation automatically.
Question 9: What does 'operational readiness review' (ORR) mean for a new on-call rotation member, and what should be verified?
- An ORR verifies that the new team member has required access, understands the services they will support, has read key runbooks, and has completed a shadow rotation before taking primary on-call responsibility (Correct answer)
- An ORR is a technical review of the new engineer's code contributions to verify they are capable of writing production-quality fixes
- An ORR is an annual certification that all on-call engineers must renew to maintain their authorization to respond to production incidents
- An ORR is a management review of on-call scheduling to ensure equal distribution of on-call hours across the team
Correct answer: An ORR verifies that the new team member has required access, understands the services they will support, has read key runbooks, and has completed a shadow rotation before taking primary on-call responsibility
An ORR for a new on-call member verifies operational readiness: access provisioned, runbooks read and understood, shadow rotation completed, and critical escalation paths known β ensuring they can effectively respond before holding primary responsibility.
Question 10: What does 'runbook drift' refer to in SRE practice?
- Runbooks written by multiple authors
- Runbooks that are rarely executed
- Runbooks that become outdated as the system changes (Correct answer)
- Runbooks stored in multiple formats
Correct answer: Runbooks that become outdated as the system changes
Runbook drift occurs when documentation is not updated alongside system changes, causing instructions to become inaccurate or misleading.
Question 11: What is 'storage tiering' in cloud cost optimization, and which data access patterns benefit MOST from it?
- Storage tiering reduces storage costs by compressing all data before storing it, with the compression ratio determining the tier
- Storage tiering replicates data across multiple storage tiers simultaneously to improve read performance through parallel access
- Storage tiering is a backup strategy that stores multiple versions of data in different geographic regions for DR purposes
- Storage tiering uses cheaper storage classes (e.g., S3 Glacier, Nearline) for infrequently accessed data while keeping frequently accessed data in faster, more expensive storage β most beneficial for large datasets with clear hot/cold access patterns (Correct answer)
Correct answer: Storage tiering uses cheaper storage classes (e.g., S3 Glacier, Nearline) for infrequently accessed data while keeping frequently accessed data in faster, more expensive storage β most beneficial for large datasets with clear hot/cold access patterns
S3 Intelligent Tiering, S3 Glacier, and similar services can reduce storage costs by 70-90% for data that doesn't need frequent access. The largest savings come from cold data like audit logs, old backups, and historical analytics data.
Question 12: What is the 'reliability hierarchy' in SRE, and what does it imply about incident priorities?
- Hardware β Network β OS β Application β User β incidents should always be debugged from the bottom layer up
- Customer satisfaction β Features β Performance β Availability β reliability priorities should follow business value
- Monitoring β Incident Response β Postmortem β Automation β Capacity Planning β each layer must be solid before the next adds value; an incident cannot be well-managed if monitoring is blind (Correct answer)
- SLA β SLO β SLI β Error Budget β each layer constrains the one below it in importance
Correct answer: Monitoring β Incident Response β Postmortem β Automation β Capacity Planning β each layer must be solid before the next adds value; an incident cannot be well-managed if monitoring is blind
The SRE reliability pyramid establishes that without solid monitoring, incident response is blind; without good incident response, postmortems lack data; without postmortems, automation lacks direction β each level depends on the foundation below it.
Question 13: What is the purpose of a 'sidecar proxy' in a service mesh architecture?
- To provide a redundant backup instance that takes over if the main service container crashes
- To act as an API gateway that routes external traffic to the correct microservices
- To handle cross-cutting concerns like load balancing, circuit breaking, observability, and mTLS on behalf of the service without modifying the service code (Correct answer)
- To cache database query results locally to reduce backend load
Correct answer: To handle cross-cutting concerns like load balancing, circuit breaking, observability, and mTLS on behalf of the service without modifying the service code
A sidecar proxy (e.g., Envoy in Istio) runs alongside each service instance and intercepts all network traffic, transparently applying retries, circuit breaking, mutual TLS, distributed tracing, and load balancing without any changes to the service code.
Question 14: Why is post-incident review important?
- To analyze the root cause and prevent recurrence (Correct answer)
- To punish the responsible person.
- To ensure compliance with HR rules.
- To find bugs in unrelated systems.
Correct answer: To analyze the root cause and prevent recurrence
Post-incident reviews, often called blameless postmortems, are crucial for learning from failures and improving system resilience. Their purpose is to thoroughly analyze the root cause of an incident, identify contributing factors, and implement preventative measures to avoid similar issues in the future. This process fosters a culture of continuous improvement rather than assigning blame.
Question 15: What is the primary purpose of a 'steady state hypothesis' in chaos engineering?
- To document the normal behavior of a system before introducing failures (Correct answer)
- To predict which components will fail during an experiment
- To establish SLO thresholds for production traffic
- To define the maximum acceptable downtime for a service
Correct answer: To document the normal behavior of a system before introducing failures
The steady state hypothesis defines observable, measurable normal behavior so you can verify the system returns to that state after chaos is injected.
Question 16: Which command initializes a Terraform working directory and downloads required providers?
- terraform install
- terraform init (Correct answer)
- terraform get
- terraform setup
Correct answer: terraform init
`terraform init` initializes the working directory, downloads provider plugins, and sets up the backend for state storage.
Question 17: Which of the following is a benefit of automation in infrastructure?
- Requires more manual oversight.
- Promotes inconsistent configurations.
- Increases manual configuration tasks.
- Reduces deployment time and human errors (Correct answer)
Correct answer: Reduces deployment time and human errors
Automation in infrastructure management significantly streamlines repetitive tasks, leading to much faster provisioning and deployment cycles. By replacing manual steps with automated scripts and tools, it drastically reduces the likelihood of human errors, ensuring greater consistency and reliability across environments. This frees up SREs to focus on more complex problem-solving and strategic initiatives rather than repetitive manual work.
Question 18: In a GitOps workflow, what serves as the single source of truth for infrastructure state?
- A Git repository (Correct answer)
- The running production environment
- A configuration management database
- The CI/CD pipeline logs
Correct answer: A Git repository
GitOps uses a Git repository as the authoritative source of truth, with automated processes reconciling actual infrastructure state to match it.
Question 19: In Prometheus, what does the `rate()` function compute?
- The per-second average increase of a counter over a specified time range (Correct answer)
- The 95th percentile latency over a time window
- The current value of a gauge metric
- The sum of all metric values across all instances
Correct answer: The per-second average increase of a counter over a specified time range
`rate()` calculates the per-second average rate of increase for a counter metric over the given range, accounting for counter resets.
Question 20: What does Amdahl's Law imply for SRE capacity planning when scaling parallel systems?
- Doubling CPU cores always halves processing time
- Performance scales linearly with added resources
- The speedup from parallelization is limited by the sequential (non-parallelizable) fraction of the workload (Correct answer)
- Network bandwidth is always the primary bottleneck
Correct answer: The speedup from parallelization is limited by the sequential (non-parallelizable) fraction of the workload
Amdahl's Law states that sequential portions of a workload cap the maximum speedup achievable through parallelization, setting an upper bound on scaling benefits.
Question 21: Which scenario BEST illustrates 'toil that creates more toil' (a toil feedback loop)?
- Investigating a novel incident not seen before
- Reviewing pull requests for the automation framework
- Manually scaling servers under load generates alerts requiring manual acknowledgment, which itself generates tickets requiring manual closure (Correct answer)
- Writing automation code that must be maintained
Correct answer: Manually scaling servers under load generates alerts requiring manual acknowledgment, which itself generates tickets requiring manual closure
Toil feedback loops occur when one manual task triggers additional manual tasks, compounding the time burden with each cycle.
Question 22: What is 'incident commander' (IC) role and what are the TOP THREE responsibilities during a major incident?
- Coordinate all responders and communications, make decisions on mitigation actions and escalations, and maintain situational awareness across the incident β NOT to personally diagnose or fix the technical issue (Correct answer)
- Write the postmortem, manage the on-call rotation, and deploy the hotfix that resolves the incident
- Contact affected customers, refund SLA credits, and brief the board of directors on the incident impact
- Monitor all metrics dashboards, approve all production changes, and write the final root cause analysis
Correct answer: Coordinate all responders and communications, make decisions on mitigation actions and escalations, and maintain situational awareness across the incident β NOT to personally diagnose or fix the technical issue
The IC's value is coordination and decision-making authority, not technical execution. By delegating technical tasks, the IC maintains the bird's-eye view needed to make tactical decisions and keep the incident moving toward resolution.
Question 23: What is the 'recovery hierarchy' concept in business continuity, and how should SREs use it to prioritize restoration order?
- The recovery hierarchy ranks services by the alphabetical order of their names for consistent and unambiguous restoration sequencing
- The recovery hierarchy is the order in which backup tapes are processed, starting with the most recent backup
- The recovery hierarchy defines which team members are paged first during a disaster, based on seniority and availability
- The recovery hierarchy ranks services by their criticality to core business function and dependencies, ensuring foundational services (authentication, networking, databases) are restored before dependent services (application tiers, reporting) (Correct answer)
Correct answer: The recovery hierarchy ranks services by their criticality to core business function and dependencies, ensuring foundational services (authentication, networking, databases) are restored before dependent services (application tiers, reporting)
Restoring services in dependency order prevents wasted effort β an application server is useless without its database; a dashboard is useless without the services it monitors. The hierarchy maps dependencies to a restoration sequence.
Question 24: What is the role of version control in change management?
- It limits collaboration among developers.
- It prevents testing from happening.
- It supports visibility and rollback during changes (Correct answer)
- It removes the need for documentation.
Correct answer: It supports visibility and rollback during changes
Version control systems (like Git) are fundamental to effective change management. They provide a complete, auditable history of all changes made to configuration files, code, and infrastructure definitions. This visibility allows teams to understand the evolution of systems and, critically, enables quick and reliable rollbacks to previous stable states if a change introduces problems, thereby minimizing downtime and supporting safe changes.
Question 25: A team uses Elasticsearch, Logstash, and Kibana (ELK stack). What is Logstash's role?
- Ingesting, transforming, and shipping log data to Elasticsearch (Correct answer)
- Providing the visualization and dashboarding interface
- Storing and indexing log data for search
- Collecting metrics from application endpoints via scraping
Correct answer: Ingesting, transforming, and shipping log data to Elasticsearch
Logstash is the data processing pipeline that ingests logs from various sources, applies filters and transformations, and forwards them to Elasticsearch.
Question 26: Which principle emphasizes balancing innovation with reliability?
- Error budgeting (Correct answer)
- Resource capping
- 100% uptime policy
- Load balancing
Correct answer: Error budgeting
Error budgeting is a core SRE principle that explicitly balances the tension between innovation (deploying new features) and reliability (maintaining system stability). By defining an acceptable level of unreliability (the error budget), teams can use the remaining budget for experimentation and new deployments. If the budget is exhausted, focus shifts back to reliability work, ensuring a sustainable pace of development and balancing these two critical aspects.
Question 27: According to the Google SRE book, which question should every alert be able to answer 'yes' to in order to justify its existence?
- Has this alert fired at least once in production in the past 90 days?
- Does this alert correlate with at least one other metric to confirm its accuracy?
- Is this alert actionableβdoes a human need to do something in response right now? (Correct answer)
- Does this alert have a corresponding automated remediation script attached?
Correct answer: Is this alert actionableβdoes a human need to do something in response right now?
The Google SRE book's core alerting principle states that every page-worthy alert must require immediate human action; alerts that don't need a human response right now should be tickets or removed entirely.
Question 28: An on-call engineer receives a P1 alert at 2 AM but cannot determine the cause after 20 minutes of investigation. What is the CORRECT action?
- Resolve the alert as 'no issue found' and create a low-priority ticket for investigation during business hours
- Escalate to the secondary on-call or the designated expert as defined in the escalation policy, without waiting longer β unresolved P1 incidents require additional resources (Correct answer)
- Attempt to mitigate by restarting all services in the affected stack, then monitor to see if the alert clears
- Continue investigating independently for another hour before escalating to avoid waking others unnecessarily
Correct answer: Escalate to the secondary on-call or the designated expert as defined in the escalation policy, without waiting longer β unresolved P1 incidents require additional resources
P1 incidents have defined escalation timeouts in the incident management policy β waiting beyond them is a policy violation that risks extended user impact. Escalation is not failure; it is the defined process.
Question 29: A reliability review finds that a service achieves 99.97% availability β well above its 99.9% SLO. The SRE team proposes loosening the SLO to 99.95%. What is the BEST argument in favor of this change?
- A looser SLO will attract fewer customer complaints because expectations are lower
- The SLO should be loosened to match the SLA to reduce confusion
- The gap between actual reliability and the SLO target suggests the service is over-engineered for its current needs; loosening the SLO frees error budget for feature velocity without meaningfully impacting users (Correct answer)
- Loosening the SLO reduces engineering accountability and makes the team's job easier
Correct answer: The gap between actual reliability and the SLO target suggests the service is over-engineered for its current needs; loosening the SLO frees error budget for feature velocity without meaningfully impacting users
If a service consistently and significantly exceeds its SLO, it may be over-engineered relative to user needs, consuming engineering resources that could be used for features. Adjusting the SLO to reflect the actual user-acceptable threshold releases that budget.
Question 30: When a SRE professional encounters an unfamiliar challenge in capacity planning & scaling, what is the recommended first course of action?
- Proceed based on personal intuition alone
- Apply the solution used for the most recent similar problem without adaptation
- Postpone addressing the issue indefinitely
- Research applicable standards, consult with subject matter experts, and document the approach (Correct answer)
Correct answer: Research applicable standards, consult with subject matter experts, and document the approach
Professional practice requires a methodical approach to unfamiliar challenges: research the applicable standards, consult experts when needed, and document the reasoning for the chosen approach.
Question 31: A microservice running in Kubernetes is experiencing intermittent OOMKilled events. What is the MOST appropriate first response?
- Increase the memory limit to 10Γ the current value to prevent future OOMKilled events
- Analyze memory usage patterns using profiling tools to identify memory leaks or excessive allocation, then set appropriate memory limits and requests based on observed usage (Correct answer)
- Remove all memory limits so Kubernetes does not kill the pod when memory spikes
- Scale the number of replicas to distribute memory load across more pods
Correct answer: Analyze memory usage patterns using profiling tools to identify memory leaks or excessive allocation, then set appropriate memory limits and requests based on observed usage
OOMKilled events indicate the container exceeded its memory limit. The correct approach is to investigate whether the limit is too low for legitimate usage or whether there is a memory leak, then set limits based on profiled actual usage.
Question 32: A team's on-call rotation handles 50 alerts per week, 80% of which are auto-resolved before engineers investigate. What should be the FIRST step?
- Silence the alerts that always auto-resolve and add automation to handle them (Correct answer)
- Hire more on-call engineers to share the load
- Increase alert thresholds to reduce noise
- Archive the alerts for monthly review
Correct answer: Silence the alerts that always auto-resolve and add automation to handle them
Alerts that consistently auto-resolve without human action are pure toil β they should be automated end-to-end or suppressed with automated remediation.
Question 33: An SRE notices that every deployment requires manual SSH into servers to update config files. What practice should they implement?
- Delegate manual config updates to a dedicated ops team
- Add a reminder step in the deployment checklist
- Configuration as code β version-control config and apply it automatically as part of the deployment pipeline (Correct answer)
- Increase the deployment frequency to reduce config drift
Correct answer: Configuration as code β version-control config and apply it automatically as part of the deployment pipeline
Configuration as code ensures config changes are versioned, reviewed, and deployed automatically alongside application changes.
Question 34: Which property of a good SLO measurement window is MOST important for catching slow reliability degradations?
- Using a very short window (e.g., 1 hour) so problems are detected as quickly as possible
- Using a rolling window (e.g., trailing 30 days) rather than a fixed calendar window, so degradation is continuously tracked (Correct answer)
- Resetting the window weekly so each week gets a fresh start
- Using a calendar quarter window to align reliability tracking with business reporting cycles
Correct answer: Using a rolling window (e.g., trailing 30 days) rather than a fixed calendar window, so degradation is continuously tracked
Rolling windows continuously reflect the current reliability trajectory. A fixed calendar window can hide a slow degradation that started late in one period and continues into the next β the calendar reset discards accumulated history.
Question 35: What does 'pipeline as code' mean in modern CI/CD practices?
- Writing pipeline logic in the application's main codebase
- Storing pipeline execution logs in the code repository
- Defining pipeline configuration in version-controlled files alongside application code (Correct answer)
- Using a GUI to configure pipeline steps
Correct answer: Defining pipeline configuration in version-controlled files alongside application code
Pipeline as code means storing pipeline definitions (e.g., Jenkinsfile, .gitlab-ci.yml) in version control for auditability and reproducibility.
Question 36: A runbook step says 'run this SQL query monthly to archive old records.' How should an SRE classify and address this?
- Classify as toil and automate with a scheduled cron job or database maintenance policy (Correct answer)
- Convert it to a quarterly task to reduce frequency
- Keep it in the runbook since it only runs once a month
- Delegate it permanently to the database team
Correct answer: Classify as toil and automate with a scheduled cron job or database maintenance policy
Repetitive manual tasks with predictable triggers are classic toil regardless of frequency and should be automated away.
Question 37: A service has been paging on-call engineers 15 times per shift on average. According to SRE principles, what should be done?
- Reduce monitoring coverage to lower page count
- Investigate and eliminate the sources of noise through alert tuning or automation (Correct answer)
- Accept the high page volume as a sign of a critical service
- Increase the on-call team size to share the load
Correct answer: Investigate and eliminate the sources of noise through alert tuning or automation
High page volume indicates poor alert hygiene; SRE practice calls for eliminating noise through tuning thresholds, improving automation, and fixing root causes.
Question 38: An alert fires at 3 AM and the on-call engineer determines no user impact exists. What is the appropriate action?
- Immediately page the entire team to investigate
- Acknowledge, document the lack of impact, and create a ticket to tune the alert (Correct answer)
- Escalate to the VP of Engineering
- Silence the alert permanently without investigation
Correct answer: Acknowledge, document the lack of impact, and create a ticket to tune the alert
Non-impacting alerts should be acknowledged and documented, then a follow-up ticket created to reduce noise β silencing without investigation risks missing real issues.
Question 39: In the context of SRE certification, what is the most important consideration when implementing distributed systems design?
- Ensuring alignment with established standards, stakeholder needs, and best practices (Correct answer)
- Minimizing documentation to save time
- Delegating all responsibilities to junior staff
- Completing implementation as quickly as possible regardless of quality
Correct answer: Ensuring alignment with established standards, stakeholder needs, and best practices
When implementing distributed systems design, SRE professionals must ensure alignment with industry standards and stakeholder needs. Hasty implementation without proper planning often leads to compliance issues and suboptimal outcomes.
Question 40: What is the most effective way to measure success in toil reduction & elimination within SRE professional practice?
- Rely solely on supervisor opinion
- Count only the number of activities completed
- Compare only with industry averages without considering context
- Use a combination of quantitative metrics, qualitative assessments, and stakeholder feedback aligned with defined objectives (Correct answer)
Correct answer: Use a combination of quantitative metrics, qualitative assessments, and stakeholder feedback aligned with defined objectives
Effective measurement combines multiple data sources β quantitative metrics, qualitative assessments, and stakeholder feedback β all aligned with clearly defined objectives for a comprehensive evaluation.
SRE Foundation Certification Exam
The SRE Foundation Certification Exam validates an individual's understanding of the core principles and practices of Site Reliability Engineering (SRE).
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong β answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds