← All DSE Flashcard Decks

Model Evaluation and Validation Flashcards

7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Model Evaluation and Validation flashcards as text
  1. What does the Brier Score measure in probabilistic classification models?

    Answer: Mean squared error between predicted probabilities and actual outcomes

    The Brier Score is the mean squared difference between predicted probabilities and binary outcomes, where lower scores indicate better-calibrated predictions.

  2. In a reliability diagram (calibration plot), a perfectly calibrated model's curve would:

    Answer: Lie along the diagonal from (0,0) to (1,1)

    Perfect calibration means that when a model predicts 70% probability, 70% of those cases are truly positive, forming a diagonal line.

  3. Which technique adjusts a classifier's predicted probabilities to be better calibrated without retraining it?

    Answer: Platt scaling

    Platt scaling fits a logistic regression on the model's raw outputs to transform them into calibrated probabilities.

  4. When using RMSE versus MAE to evaluate regression models, RMSE is preferred when:

    Answer: Large errors are particularly undesirable

    RMSE squares errors before averaging, so it penalizes large errors more heavily than MAE, making it preferable when large deviations are costly.

  5. What is the key difference between R² (coefficient of determination) and Adjusted R²?

    Answer: Adjusted R² penalizes for the number of predictors added

    Adjusted R² penalizes for adding irrelevant predictors, unlike R² which always increases or stays the same as features are added.

  6. In nested cross-validation, the inner loop is used for:

    Answer: Hyperparameter selection

    The inner loop performs hyperparameter tuning, while the outer loop provides an unbiased estimate of the tuned model's generalization error.

  7. Which situation best describes when leave-one-out cross-validation (LOOCV) is most justified?

    Answer: Very small datasets where every sample matters

    LOOCV is computationally expensive but maximally uses available data, making it most valuable when datasets are too small for larger holdout folds.