Model Evaluation and Validation Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Model Evaluation and Validation flashcards as text
Which of the following is a consequence of using the test set multiple times during model development?
Answer: Optimistic bias in generalization estimates
Repeatedly evaluating on the test set allows implicit overfitting to it, making the reported metrics overly optimistic and no longer representative of true generalization.
Lift in model evaluation is calculated as:
Answer: (Precision of model) / (Precision of random model)
Lift measures how much better the model performs compared to a random baseline, calculated as the model's response rate in a decile divided by the overall response rate.
What is the purpose of a reliability diagram (calibration plot)?
Answer: To visualize whether predicted probabilities match observed event frequencies
A reliability diagram bins predictions by predicted probability and plots mean predicted probability versus observed frequency, revealing over- or under-confidence.
In a multiclass classification problem, macro-averaged F1 differs from weighted-averaged F1 in that macro-averaging:
Answer: Gives equal weight to each class regardless of class frequency
Macro-averaging computes F1 for each class independently and takes an unweighted mean, treating all classes equally regardless of how many samples they contain.
What does a learning curve showing high training error and high validation error indicate?
Answer: Underfitting — model lacks capacity to capture the pattern
When both training and validation errors are high and converge, the model is underfitting because it cannot capture the underlying pattern in the data.
Platt scaling is a technique used to:
Answer: Convert a classifier's raw scores into calibrated probability estimates
Platt scaling fits a logistic regression on the classifier's output scores to transform them into well-calibrated posterior probabilities.
When comparing multiple models across several datasets, the Friedman test followed by a Nemenyi post-hoc test is preferred over pairwise t-tests because:
Answer: It controls the family-wise error rate inflated by multiple comparisons
The Friedman-Nemenyi procedure is a non-parametric approach that corrects for multiple comparisons, preventing inflated Type I error rates when simultaneously comparing many models.