AI Engineer: MLOps and Model Deployment Flashcards
6 cards from real AI practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 AI Engineer: MLOps and Model Deployment flashcards as text
What is the purpose of 'model quantization' in the context of ML deployment?
Answer: To reduce model size and inference time by using lower-precision arithmetic
Quantization converts model weights from 32-bit floats to 8-bit integers, shrinking memory footprint and speeding up inference with minimal accuracy loss.
Which Python library is widely used for building ML pipelines that can be tracked and reproduced?
Answer: DVC (Data Version Control)
DVC tracks data, models, and pipeline stages in Git-compatible workflows, enabling reproducible ML experiments.
What does 'SLA' stand for in the context of ML model serving?
Answer: Service Level Agreement
A Service Level Agreement defines the agreed-upon performance targets, such as maximum latency and uptime, for a model serving system.
In MLOps, what does 'continuous training' (CT) refer to?
Answer: Automatically retraining models on new data triggered by schedule or drift detection
Continuous training automatically triggers model retraining on new production data, often scheduled or triggered by detected data or concept drift.
Which serving framework is designed for high-performance, production-grade deployment of ML models, supporting REST and gRPC?
Answer: TensorFlow Serving
TensorFlow Serving is a flexible, high-performance serving system for ML models, supporting REST and gRPC endpoints for production deployment.
What is the main benefit of using a directed acyclic graph (DAG) to represent an ML pipeline?
Answer: It explicitly defines task dependencies, enabling parallel execution and reproducibility
A DAG maps task dependencies so the orchestrator can parallelize independent steps and ensure correct execution order for reproducible pipelines.