Model Evaluation and Validation Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Model Evaluation and Validation flashcards as text
When performing nested cross-validation, the outer loop is used for:
Answer: Unbiased generalization error estimation
Nested CV uses an inner loop for hyperparameter tuning and model selection, while the outer loop provides an unbiased estimate of the final model's generalization error.
A regression model reports an R² of 0.85 on training data and 0.42 on test data. What is the most likely diagnosis?
Answer: Overfitting
A large gap between training R² and test R² indicates the model memorized training data patterns that do not generalize, a hallmark of overfitting.
Which of the following is the correct interpretation of a 95% confidence interval for model accuracy?
Answer: If the experiment were repeated many times, 95% of such intervals would contain the true accuracy
A confidence interval is a frequentist concept meaning that the procedure generates intervals containing the true parameter 95% of the time across repeated experiments.
What does the McNemar test evaluate in model comparison?
Answer: Whether the disagreements between two classifiers are statistically symmetric
McNemar's test uses a 2x2 contingency table of cases where classifiers disagree to test if one model significantly outperforms the other.
In time-series model validation, why is standard k-fold cross-validation inappropriate?
Answer: It can cause temporal leakage by using future data to predict the past
Random splitting in k-fold CV allows future observations to appear in the training set, creating data leakage that violates the temporal structure of time-series data.
A confusion matrix shows TP=80, FP=20, FN=10, TN=90. What is the F1 score?
Answer: 0.84
Precision = 80/100 = 0.80, Recall = 80/90 ≈ 0.889; F1 = 2*(0.80*0.889)/(0.80+0.889) ≈ 0.842.
What is the primary purpose of using a holdout validation set separate from both training and test sets?
Answer: To tune hyperparameters without contaminating the final test evaluation
A separate validation set allows hyperparameter tuning while preserving the test set as a truly unseen benchmark for final model evaluation.