โ† All Data Science Flashcard Decks

Data Science Supervised Learning Models Questions and Answers Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Data Science Supervised Learning Models Questions and Answers flashcards as text
  1. Which supervised learning algorithm constructs a series of if-then rules by recursively partitioning the feature space?

    Answer: Decision Tree

    Decision trees split data recursively using feature thresholds to create interpretable if-then decision rules.

  2. What is the primary purpose of regularization in supervised learning models such as Ridge and Lasso regression?

    Answer: To prevent overfitting by penalizing large coefficients

    Regularization adds a penalty term to the loss function that discourages overly complex models with large coefficient values.

  3. In a Random Forest classifier, what technique is used to reduce variance and improve generalization?

    Answer: Bagging with random feature subsets

    Random Forest combines bagging (bootstrap aggregating) with random feature selection at each split to decorrelate trees and reduce variance.

  4. Which metric is most appropriate for evaluating a supervised classification model when the dataset has a severe class imbalance?

    Answer: Area Under the Precision-Recall Curve (AUPRC)

    AUPRC focuses on the performance for the minority class and is more informative than accuracy when classes are heavily imbalanced.

  5. What does the kernel trick enable Support Vector Machines to do?

    Answer: Map data into a higher-dimensional space to find nonlinear decision boundaries

    The kernel trick implicitly maps input features into a higher-dimensional space where a linear separator can capture nonlinear relationships.

  6. In gradient boosting, what does each successive tree in the ensemble attempt to model?

    Answer: The residual errors of the previous ensemble

    Each new tree in gradient boosting is trained to predict the residual errors left by the combined predictions of all preceding trees.