โ† All Data Science Flashcard Decks

Data Science Model Performance and Evaluation Questions and Answers Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Data Science Model Performance and Evaluation Questions and Answers flashcards as text
  1. When comparing two models using ROC curves, what does a curve closer to the top-left corner indicate?

    Answer: Better trade-off between sensitivity and specificity

    A ROC curve closer to the top-left corner represents high true positive rates with low false positive rates, indicating superior discriminative ability.

  2. What is the main advantage of using the F-beta score with beta=2 instead of the standard F1 score?

    Answer: It places more emphasis on recall than precision

    F-beta with beta=2 weighs recall twice as heavily as precision, making it ideal when missing positive cases is costlier than false alarms.

  3. What problem does stratified sampling in cross-validation specifically address?

    Answer: Preserving class distribution across folds

    Stratified sampling ensures each fold maintains the same proportion of each class as the full dataset, preventing biased evaluation from uneven splits.

  4. Which metric would you use to evaluate a regression model's performance in terms of the same units as the target variable?

    Answer: Root Mean Squared Error

    RMSE takes the square root of MSE, returning the error to the original units of the target variable for more interpretable results.

  5. What does a learning curve showing training and validation scores converging at a low value suggest?

    Answer: The model is underfitting and more data will not help

    When both curves converge at a low score, the model lacks capacity to learn the underlying pattern, indicating high bias that more data cannot fix.

  6. In a confusion matrix for a binary classifier, what does the false negative rate (FNR) represent?

    Answer: The proportion of actual positives that were incorrectly classified as negative

    The false negative rate is the fraction of truly positive instances that the model failed to identify, calculated as FN divided by (FN + TP).