โ† All Data Science Flashcard Decks

FREE Data Science Model Evaluation and Validation Questions and Answers Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 FREE Data Science Model Evaluation and Validation Questions and Answers flashcards as text
  1. Which metric is most appropriate for evaluating a classification model on a highly imbalanced dataset?

    Answer: Precision-Recall AUC

    Precision-Recall AUC is preferred for imbalanced datasets because accuracy can be misleadingly high when the majority class dominates predictions.

  2. What does a high variance and low bias in a model typically indicate?

    Answer: Overfitting

    High variance with low bias means the model fits training data very well but fails to generalize, which is the hallmark of overfitting.

  3. In stratified k-fold cross-validation, what is preserved across each fold?

    Answer: The proportion of each class label

    Stratified k-fold ensures each fold maintains the same class distribution as the original dataset, which is critical for representative evaluation.

  4. What is the primary purpose of a calibration curve (reliability diagram)?

    Answer: To assess whether predicted probabilities match actual outcomes

    A calibration curve plots predicted probabilities against observed frequencies to show whether a model's confidence scores are trustworthy.

  5. When using bootstrapping for model evaluation, what is the typical approach?

    Answer: Sampling with replacement to create multiple training sets

    Bootstrapping creates multiple resampled datasets by sampling with replacement from the original data, allowing robust estimation of model performance statistics.

  6. What does the Kolmogorov-Smirnov (KS) statistic measure in model evaluation?

    Answer: The maximum separation between cumulative distributions of positive and negative classes

    The KS statistic measures the maximum distance between the cumulative distribution functions of positive and negative class scores, indicating discriminative power.