Quality Assurance & Improvement Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Quality Assurance & Improvement flashcards as text
In a continuous integration pipeline for ML models, a 'regression gate' is best defined as:
Answer: An automated check that blocks deployment if a new model underperforms the current production model on a reference dataset
A regression gate enforces a minimum quality bar by comparing a candidate model against the current champion before allowing promotion to production.
Which data versioning practice is most critical for ensuring reproducibility in ML QA?
Answer: Pinning exact dataset versions and transformation hashes alongside model artifacts
Reproducible ML requires the exact training data (including transformations) to be versioned alongside the model, since code and hyperparameters alone cannot reconstruct an experiment.
When comparing two ML models using A/B testing in production, the minimum required sample size is determined by:
Answer: Desired statistical power, significance level, and minimum detectable effect size
Sample size for A/B tests is calculated from the power (1-β), significance threshold (α), and the smallest effect the test must reliably detect.
Mutation testing in ML software QA involves:
Answer: Introducing deliberate bugs into code and verifying that tests catch them
Mutation testing assesses test suite quality by checking whether intentionally introduced code faults cause test failures, revealing gaps in test coverage.
In MLOps, 'model lineage' tracking primarily enables which QA capability?
Answer: Tracing any deployed model back to its exact training data, code, and environment
Model lineage records the provenance chain—data, code, dependencies, and parameters—enabling full auditability and root-cause analysis when issues arise.
Which statistical method is recommended for comparing the performance of two classifiers across multiple datasets to avoid relying on a single test set?
Answer: Wilcoxon signed-rank test on per-dataset metric differences
The Wilcoxon signed-rank test is a non-parametric alternative to the paired t-test that is more robust when performance differences are not normally distributed across datasets.
What is the primary purpose of 'stress testing' a machine learning model as part of QA?
Answer: Evaluating model behavior under extreme, rare, or adversarial input conditions beyond the training distribution
Stress testing exposes failure modes by deliberately presenting inputs that challenge the model's assumptions, revealing brittleness before production deployment.