Data Wrangling and Preprocessing Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Wrangling and Preprocessing flashcards as text
A feature ranges from 0 to 1,000,000 while another ranges from 0 to 1. Which preprocessing step puts them on a comparable scale?
Answer: Feature scaling (normalization/standardization)
Scaling transforms features to comparable ranges so models relying on distance aren't dominated by large-magnitude features.
Min-max scaling transforms a feature to which range by default?
Answer: [0, 1]
Min-max scaling maps the minimum to 0 and the maximum to 1 using (x - min)/(max - min).
Standardization (z-score) rescales a feature to have what properties?
Answer: Mean 0 and standard deviation 1
Z-score standardization subtracts the mean and divides by the standard deviation, yielding mean 0 and unit variance.
When encoding a nominal categorical variable with no inherent order, which method avoids implying a false ranking?
Answer: One-hot encoding
One-hot encoding creates a binary column per category, avoiding the implied ordering that integer label encoding introduces.
A right-skewed feature like website session duration is transformed with a log. What is the main benefit?
Answer: It reduces skewness and compresses large values
A log transform compresses the long right tail, reducing skewness and lessening the influence of extreme values.
Why should scaling parameters be computed only on the training set, then applied to the test set?
Answer: To prevent data leakage from the test set
Fitting the scaler on training data only prevents test-set statistics from leaking into the model and inflating performance.
Which scaling method is most robust when a feature contains many extreme outliers?
Answer: RobustScaler (uses median and IQR)
RobustScaler centers on the median and scales by the interquartile range, making it resistant to outliers.