Supervised Learning Algorithms Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Supervised Learning Algorithms flashcards as text
Which of the following is a key difference between bagging and boosting ensemble methods?
Answer: Bagging trains models in parallel on bootstrapped samples; boosting trains sequentially correcting prior errors
Bagging builds independent models in parallel on bootstrap samples and averages results, while boosting builds models sequentially where each corrects the errors of its predecessor.
In a neural network, what does batch normalization do during training?
Answer: It normalizes activations of each layer across the mini-batch, stabilizing training
Batch normalization standardizes activations within each mini-batch, reducing internal covariate shift and allowing higher learning rates and faster convergence.
What is the 'curse of dimensionality' and how does it affect k-NN classifiers?
Answer: In high dimensions, data becomes sparse and distance metrics lose discriminative power
As dimensionality grows, data points become increasingly equidistant, making nearest-neighbor distances meaningless and degrading k-NN performance significantly.
Which of the following describes the purpose of cross-validation in supervised learning?
Answer: To estimate a model's generalization performance and tune hyperparameters on limited data
Cross-validation partitions data into multiple folds, training and evaluating the model on different subsets to produce a reliable estimate of generalization error.
A decision tree trained to zero training error likely suffers from which problem?
Answer: Overfitting due to memorizing training noise
A tree that perfectly fits training data has memorized noise and specific examples, resulting in high variance and poor generalization to unseen data.
In the context of AdaBoost, how are training samples reweighted across iterations?
Answer: Misclassified samples receive higher weights so subsequent classifiers focus on hard examples
AdaBoost increases the weights of misclassified samples after each round, forcing the next weak learner to concentrate on the examples the current ensemble handles poorly.
What is the main reason Lasso (L1) regression can produce sparse models while Ridge (L2) typically does not?
Answer: The L1 penalty creates corners in the constraint region where coefficients are exactly zero
The diamond-shaped L1 constraint region has corners aligned with the axes; solutions are likely to occur at these corners where some coefficients are exactly zero, producing sparsity.