Project Planning & Execution Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Project Planning & Execution flashcards as text
In an ML project using the CRISP-DM methodology, which phase produces the 'data quality report'?
Answer: Data Understanding
The Data Understanding phase in CRISP-DM includes exploratory data analysis and producing a data quality report identifying missing values, outliers, and distribution issues.
What is the MAIN advantage of maintaining separate training, validation, and test splits that are immutable throughout the project?
Answer: It prevents data leakage and ensures the test set remains an unbiased estimate of real-world generalization
Immutable splits prevent inadvertent leakage from hyperparameter tuning decisions bleeding into the test set, preserving its role as an unbiased final evaluation.
A machine learning team uses continuous integration (CI) for model code. Which check is MOST valuable to include in the ML CI pipeline beyond standard unit tests?
Answer: Training a fast smoke-test model run to verify the pipeline executes end-to-end and metrics are within expected bounds
An ML-specific CI check that runs a mini training job catches broken pipelines, metric regressions, and data schema changes before they reach production.
When a model is promoted from staging to production in a blue-green deployment, what does 'green' represent?
Answer: The new model version that receives traffic once validated, while the old version remains on standby
In blue-green deployment, green is the new version brought up alongside the existing blue version; traffic is switched to green after validation, with blue available for instant rollback.
What does a 'definition of done' for an ML task typically include that differs from a standard software definition of done?
Answer: Performance thresholds on held-out data, bias checks across demographic subgroups, and documented failure modes
ML tasks require quantitative performance gates, fairness checks, and documented failure modes as part of the definition of done, in addition to standard code quality criteria.
Which practice BEST mitigates the risk of training-serving skew in production ML systems?
Answer: Using the same preprocessing code path for both training and online inference, typically via shared feature stores or serialized transformers
Training-serving skew arises when preprocessing differs between training and inference; using identical code paths or serialized transformation objects eliminates this discrepancy.
In ML project planning, 'shadow mode deployment' is used to:
Answer: Route real production traffic to a new model without exposing its outputs to users, enabling comparison against the incumbent without user impact
Shadow mode runs the new model in parallel on live traffic, logging predictions for offline comparison against the current system without affecting actual user-facing decisions.