โ† All MS-DS Master of Data science Flashcard Decks

Model Evaluation and Validation Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Model Evaluation and Validation flashcards as text
  1. Which evaluation metric is most appropriate when the cost of a false negative is much higher than the cost of a false positive?

    Answer: Recall

    Recall (sensitivity) measures the proportion of actual positives correctly identified, making it critical when missing a positive case (false negative) is very costly.

  2. In k-fold cross-validation, what happens as k increases toward n (leave-one-out CV)?

    Answer: Bias decreases and variance increases

    As k approaches n, each fold uses nearly all data for training (low bias), but the estimates become highly variable across folds (high variance).

  3. The Brier score is used to evaluate which type of model output?

    Answer: Probability estimates from a probabilistic classifier

    The Brier score measures the mean squared difference between predicted probabilities and actual binary outcomes, assessing probabilistic calibration.

  4. A model achieves 99% accuracy on a dataset where 99% of examples are negative. This is an example of:

    Answer: The accuracy paradox

    The accuracy paradox occurs when high accuracy is misleading because a naive model predicting the majority class always scores well on imbalanced datasets.

  5. What does the area under the Precision-Recall curve (AUC-PR) measure?

    Answer: The overall quality of ranked probability predictions for imbalanced datasets

    AUC-PR summarizes the precision-recall tradeoff across all thresholds and is preferred over AUC-ROC when classes are highly imbalanced.

  6. Which technique specifically tests whether a model has learned spurious correlations by evaluating it on data where the target variable is shuffled?

    Answer: Permutation test

    A permutation test shuffles the labels and measures model performance; if the model scores high on shuffled data, it likely learned noise rather than true signal.

  7. Expected Calibration Error (ECE) measures:

    Answer: The gap between a model's confidence and its actual accuracy

    ECE quantifies how well a model's predicted probabilities align with empirical frequencies, with lower ECE indicating better-calibrated predictions.