Machine Learning & Predictive Analytics Flashcards
7 cards from real DAC practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Machine Learning & Predictive Analytics flashcards as text
Which of the following is an unsupervised learning task?
Answer: Grouping customers into segments without labels
Clustering customers without predefined labels is a classic unsupervised learning task.
What does regularization (L1 or L2) primarily do in a regression model?
Answer: Penalizes large coefficients to reduce overfitting
Regularization adds a penalty on coefficient size, discouraging overly complex models and reducing overfitting.
L1 regularization (Lasso) is especially useful because it can do what?
Answer: Drive some coefficients exactly to zero for feature selection
Lasso can shrink some coefficients to exactly zero, effectively performing automatic feature selection.
In k-means clustering, how is the number of clusters typically chosen?
Answer: Using the elbow method on within-cluster variance
The elbow method plots within-cluster variance against k and picks the point of diminishing returns.
What problem does principal component analysis (PCA) primarily address?
Answer: Dimensionality reduction
PCA reduces the number of features by projecting data onto directions of maximum variance.
A model shows high error on both training and test sets. What is the most likely issue?
Answer: Underfitting
High error on both sets indicates the model is too simple to capture the underlying pattern, i.e., underfitting.
Which approach best prevents data leakage when scaling features?
Answer: Fit the scaler only on training data, then apply to test data
Fitting the scaler only on training data prevents test-set information from leaking into the model.