โ† All DSE Flashcard Decks

Supervised Learning: Classification Flashcards

7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Supervised Learning: Classification flashcards as text
  1. In a decision tree, what criterion does the Gini impurity measure?

    Answer: The probability of misclassifying a randomly chosen element

    Gini impurity measures the probability that a randomly selected sample would be incorrectly classified if labeled according to the class distribution in the node.

  2. What is the effect of setting a very high value for the regularization parameter C in an SVM?

    Answer: Narrower margin with fewer misclassifications allowed

    A high C value penalizes misclassification heavily, leading to a narrower margin and a model that tries hard to correctly classify all training points (risk of overfitting).

  3. Which of the following is an assumption of Gaussian Naive Bayes?

    Answer: Features follow a Gaussian distribution within each class

    Gaussian Naive Bayes assumes continuous features are normally distributed within each class.

  4. When using k-Nearest Neighbors for classification, what is the effect of choosing a very large k?

    Answer: The decision boundary becomes smoother and simpler

    Large k values smooth out the decision boundary by averaging over more neighbors, reducing variance but potentially increasing bias.

  5. The ROC curve plots which two quantities against each other?

    Answer: True Positive Rate vs. False Positive Rate

    The ROC curve plots the True Positive Rate (sensitivity) on the y-axis against the False Positive Rate (1 - specificity) on the x-axis.

  6. Which of the following best describes a multiclass classification strategy known as 'One-vs-Rest' (OvR)?

    Answer: One binary classifier is trained per class, treating that class as positive and all others as negative

    In OvR, one binary classifier is trained per class to distinguish it from all other classes combined.

  7. A classifier achieves 99% accuracy on a dataset where 99% of samples belong to class 0. What does this most likely indicate?

    Answer: The model may be simply predicting class 0 for every sample

    On a highly imbalanced dataset, a trivial classifier that always predicts the majority class can achieve high accuracy, making accuracy misleading.