Cloud Infrastructure Flashcards
7 cards from real CAIC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Cloud Infrastructure flashcards as text
What does 'infrastructure as code' (IaC) provide for AI platform teams?
Answer: Reproducible, version-controlled environment provisioning
IaC tools like Terraform and Pulumi let teams define cloud resources in code, enabling repeatable, auditable, and diff-able infrastructure changes.
A CAIC consultant recommends separating AI training and inference workloads into different node pools. What is the PRIMARY reason?
Answer: Training and inference require different hardware profiles and scaling behaviors
Training is batch-oriented and GPU-heavy, while inference is latency-sensitive and may need different GPU types or even CPU-only nodes, requiring distinct pools.
Which networking concept reduces AI data transfer costs when moving large datasets between cloud storage and compute within the same region?
Answer: Private endpoints / VPC peering
Private endpoints and VPC peering keep traffic on the provider's backbone, avoiding public internet egress fees and reducing latency.
What is the purpose of a feature store in cloud AI infrastructure?
Answer: To centralize, share, and serve ML features consistently across training and serving
A feature store (e.g., Feast, Tecton) ensures the same feature logic is used in training and inference, preventing training-serving skew.
An organization needs to run AI workloads that require physical hardware control and data residency, but also want elastic capacity. Which pattern fits best?
Answer: Hybrid cloud with cloud bursting
Hybrid cloud with cloud bursting keeps sensitive workloads on-premises while automatically routing overflow to the public cloud during peak demand.
Which metric is most important when evaluating cloud GPU instance performance for distributed AI training?
Answer: Inter-GPU bandwidth (e.g., NVLink or InfiniBand throughput)
Distributed training relies on frequent gradient synchronization between GPUs; high inter-GPU bandwidth (NVLink, InfiniBand) directly reduces communication bottlenecks.
What is the role of a service mesh (e.g., Istio) in a cloud-native AI inference platform?
Answer: Managing secure, observable, and policy-driven service-to-service communication
A service mesh provides mTLS encryption, traffic policies, retries, and distributed tracing between microservices including inference endpoints.