← All CAIC Flashcard Decks

System Architecture & Design Flashcards

7 cards from real CAIC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 System Architecture & Design flashcards as text
  1. In an AI system that requires real-time inference with latency under 10ms, which deployment pattern is most appropriate?

    Answer: Edge inference with model quantization

    Edge inference with quantized models minimizes network round-trip latency and enables sub-10ms response times at the point of need.

  2. Which architectural pattern best decouples AI model updates from the application logic that consumes predictions?

    Answer: Model-as-a-Service (MaaS) with versioned API contracts

    Model-as-a-Service with versioned APIs allows models to be retrained and deployed independently without changing consumer application code.

  3. A feature store in an ML system serves what primary architectural purpose?

    Answer: Providing consistent, reusable feature computations across training and serving

    A feature store eliminates training-serving skew by ensuring the same feature transformations are applied consistently during both model training and online inference.

  4. When designing an AI system for a regulated industry, which architectural component directly addresses explainability requirements?

    Answer: Model explanation layer with SHAP or LIME integration

    Integrating post-hoc explainability tools like SHAP or LIME allows the system to generate human-interpretable rationales for individual predictions required by regulators.

  5. What is the key advantage of using a shadow deployment strategy when releasing a new AI model?

    Answer: It allows production traffic comparison without exposing users to the new model's outputs

    Shadow deployment runs the new model in parallel on real traffic and logs its outputs for comparison without affecting the user experience, enabling safe pre-production validation.

  6. In a Lambda architecture for AI, the batch layer primarily serves which function?

    Answer: Reprocessing all historical data to produce accurate, comprehensive views

    The batch layer processes the entire historical dataset on a schedule to produce highly accurate, complete views that correct any errors in real-time processing.

  7. Which technique is specifically designed to prevent training-serving skew in production ML systems?

    Answer: Sharing identical feature transformation code or pipelines between training and inference

    Training-serving skew is caused by inconsistent data preprocessing; sharing the exact same transformation pipeline code between both phases eliminates this discrepancy.