Cost Optimization and Cloud Resource Management Flashcards
6 cards from real SRE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Cost Optimization and Cloud Resource Management flashcards as text
What is 'storage tiering' in cloud cost optimization, and which data access patterns benefit MOST from it?
Answer: Storage tiering uses cheaper storage classes (e.g., S3 Glacier, Nearline) for infrequently accessed data while keeping frequently accessed data in faster, more expensive storage — most beneficial for large datasets with clear hot/cold access patterns
S3 Intelligent Tiering, S3 Glacier, and similar services can reduce storage costs by 70-90% for data that doesn't need frequent access. The largest savings come from cold data like audit logs, old backups, and historical analytics data.
What is 'autoscaling over-provisioning,' and how should SREs configure scaling policies to avoid it while maintaining reliability?
Answer: Autoscaling over-provisioning occurs when scaling policies maintain too many instances during low-traffic periods; it is reduced by tuning scale-down policies (cooldown periods, step scaling) while maintaining minimum replicas sufficient for baseline SLO compliance
Scale-down policies should aggressively remove excess capacity after traffic reduces, while scale-up policies should quickly add capacity before SLOs degrade — the balance is faster scale-up, slower scale-down, with minimum replicas covering baseline traffic.
What is 'FinOps' and how does it relate to SRE practice?
Answer: FinOps is a cross-functional practice that brings financial accountability to cloud spending through collaboration between engineering, finance, and business teams; SREs contribute cloud efficiency expertise while FinOps ensures cost decisions account for reliability requirements
FinOps creates shared accountability for cloud costs across engineering, product, and finance. SREs are key contributors because they understand the relationship between cost and reliability — they can distinguish necessary reliability overhead from waste.
What is 'compute cost attribution' per request, and why is it valuable for product and SRE teams?
Answer: Computing the infrastructure cost per API request (cost = total compute cost / total requests) enables product teams to make informed decisions about feature economics and SREs to identify expensive API endpoints worth optimizing
When you know that endpoint A costs $0.001 per request and endpoint B costs $0.50 per request, you can make informed decisions: is the value delivered by B worth 500× more cost? Should it be optimized first? Should usage be restricted?
What is 'cloud waste from over-provisioned Kubernetes resources' and how is it detected and reduced?
Answer: Pods with resource requests significantly above their actual consumption waste cluster capacity by reserving CPU and memory that goes unused; detected by comparing requested vs. actual usage via metrics, and reduced by Vertical Pod Autoscaler (VPA) or manual request right-sizing
Pod resource requests determine cluster node provisioning. If pods request 4 vCPU but average 0.2 vCPU, nodes are provisioned for the 4 vCPU request — wasting 95% of allocated (and paid-for) compute.
What is 'data transfer cost optimization' in cloud environments, and which architectures minimize data egress charges?
Answer: Cloud providers charge significantly for data transferred out of their network (egress); architectures that process data close to where it is stored (locality principle), use CDNs to cache at the edge, and keep inter-service traffic within the same region and AZ minimize these charges
Cloud data egress fees (typically $0.08-0.09/GB for AWS) can be significant for data-intensive applications. Processing data in the same region as storage, using same-AZ traffic where possible, and caching at the edge with CDNs all reduce egress charges.