Master of Data Science Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Master of Data Science flashcards as text
Which regularization technique adds the sum of absolute values of coefficients as a penalty term to the loss function?
Answer: Lasso (L1)
Lasso (L1) regularization penalizes the sum of absolute values of coefficients, which can drive some coefficients exactly to zero, performing feature selection.
In the context of Bayesian inference, what does the prior distribution represent?
Answer: Beliefs about parameters before observing data
The prior distribution encodes beliefs or knowledge about model parameters before any data is observed.
Which SQL window function returns the rank of a row within a partition, with no gaps in ranking values?
Answer: DENSE_RANK()
DENSE_RANK() assigns consecutive ranks without gaps, whereas RANK() skips values when ties occur.
What is the time complexity of training a k-Nearest Neighbors classifier on n training samples with d features?
Answer: O(1) training, O(nd) at prediction
kNN is a lazy learner — training is O(1) since it just stores data, but each prediction requires computing distances to all n points across d features.
In a neural network, what problem occurs when gradients become exponentially small as they propagate backward through many layers?
Answer: Vanishing gradient
The vanishing gradient problem causes gradients to shrink exponentially during backpropagation, making it difficult to train deep networks.
Which metric is most appropriate for evaluating a classification model when the dataset has severe class imbalance?
Answer: F1-Score
F1-Score balances precision and recall, making it far more informative than accuracy when one class vastly outnumbers another.
What does the 'curse of dimensionality' primarily refer to in machine learning?
Answer: The exponential growth of data needed as feature dimensions increase
As dimensionality increases, the volume of space grows so rapidly that data becomes sparse, making distance-based methods and density estimation unreliable.