FREE Data Science Supervised Learning Algorithms Questions and Answers Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 FREE Data Science Supervised Learning Algorithms Questions and Answers flashcards as text
What is the main disadvantage of K-Nearest Neighbors (KNN) when applied to large datasets?
Answer: High computational cost at prediction time because it must compute distances to all training points
KNN is a lazy learner that stores the entire training set and computes distances to all points during prediction, making it computationally expensive for large datasets.
Which metric does the CART algorithm use by default for regression tree splits?
Answer: Mean Squared Error (MSE)
CART uses Mean Squared Error to evaluate regression splits, choosing the split that minimizes the variance of the target variable within each resulting node.
In logistic regression, what function maps the linear combination of inputs to a probability between 0 and 1?
Answer: Sigmoid function
The sigmoid (logistic) function transforms any real-valued number into a value between 0 and 1, making it suitable for probability estimation.
What happens when the number of estimators in a Random Forest is increased significantly?
Answer: Variance decreases while bias remains roughly constant, with diminishing returns
Adding more trees to a Random Forest reduces variance through averaging but does not significantly affect bias, and the improvement plateaus after a sufficient number of trees.
Which supervised learning algorithm is most appropriate when you need a probabilistic interpretation of class membership and feature importance with minimal tuning?
Answer: Logistic Regression
Logistic regression naturally outputs calibrated probabilities, provides interpretable coefficients indicating feature importance, and requires relatively little hyperparameter tuning.
What is the purpose of pruning in decision tree algorithms?
Answer: To reduce overfitting by removing branches that provide little predictive power
Pruning removes tree branches that capture noise rather than true patterns, reducing overfitting and improving generalization to unseen data.