Knowledge Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Knowledge flashcards as text
What does the term 'data leakage' mean in a machine learning pipeline?
Answer: Information from the test set or future data inadvertently influencing model training
Data leakage occurs when information outside the legitimate training data is used to build the model, producing overly optimistic evaluation results.
In the context of natural language processing, what does TF-IDF measure?
Answer: The importance of a word to a document relative to a corpus
TF-IDF weights a word by how frequently it appears in a document (TF) offset by how common it is across all documents (IDF), highlighting distinctive terms.
What is a key difference between bagging and boosting ensemble methods?
Answer: Bagging trains models in parallel on bootstrap samples; boosting trains sequentially with error-focused reweighting
Bagging trains independent learners in parallel on random data subsets to reduce variance, while boosting trains sequentially to reduce both bias and variance.
Which distance metric computes the sum of absolute differences between coordinates?
Answer: Manhattan distance
Manhattan (L1) distance sums the absolute differences along each dimension, geometrically resembling travel along city blocks.
In time series analysis, what does 'stationarity' mean?
Answer: The statistical properties (mean, variance, autocorrelation) do not change over time
A stationary time series has constant statistical properties across time, which is a prerequisite for many forecasting models such as ARIMA.
What is the purpose of one-hot encoding in machine learning preprocessing?
Answer: To convert categorical variables into binary indicator columns for use in algorithms requiring numerical input
One-hot encoding represents each category as a separate binary column, allowing algorithms that require numerical input to use categorical variables.
Which evaluation metric measures the average squared difference between predicted and actual values?
Answer: Mean Squared Error (MSE)
Mean Squared Error averages the squared prediction errors, penalizing large errors more heavily than MAE due to squaring.