← All CAIC Flashcard Decks

Cloud Infrastructure Flashcards

7 cards from real CAIC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Cloud Infrastructure flashcards as text
  1. What does 'infrastructure as code' (IaC) provide for AI platform teams?

    Answer: Reproducible, version-controlled environment provisioning

    IaC tools like Terraform and Pulumi let teams define cloud resources in code, enabling repeatable, auditable, and diff-able infrastructure changes.

  2. A CAIC consultant recommends separating AI training and inference workloads into different node pools. What is the PRIMARY reason?

    Answer: Training and inference require different hardware profiles and scaling behaviors

    Training is batch-oriented and GPU-heavy, while inference is latency-sensitive and may need different GPU types or even CPU-only nodes, requiring distinct pools.

  3. Which networking concept reduces AI data transfer costs when moving large datasets between cloud storage and compute within the same region?

    Answer: Private endpoints / VPC peering

    Private endpoints and VPC peering keep traffic on the provider's backbone, avoiding public internet egress fees and reducing latency.

  4. What is the purpose of a feature store in cloud AI infrastructure?

    Answer: To centralize, share, and serve ML features consistently across training and serving

    A feature store (e.g., Feast, Tecton) ensures the same feature logic is used in training and inference, preventing training-serving skew.

  5. An organization needs to run AI workloads that require physical hardware control and data residency, but also want elastic capacity. Which pattern fits best?

    Answer: Hybrid cloud with cloud bursting

    Hybrid cloud with cloud bursting keeps sensitive workloads on-premises while automatically routing overflow to the public cloud during peak demand.

  6. Which metric is most important when evaluating cloud GPU instance performance for distributed AI training?

    Answer: Inter-GPU bandwidth (e.g., NVLink or InfiniBand throughput)

    Distributed training relies on frequent gradient synchronization between GPUs; high inter-GPU bandwidth (NVLink, InfiniBand) directly reduces communication bottlenecks.

  7. What is the role of a service mesh (e.g., Istio) in a cloud-native AI inference platform?

    Answer: Managing secure, observable, and policy-driven service-to-service communication

    A service mesh provides mTLS encryption, traffic policies, retries, and distributed tracing between microservices including inference endpoints.