FREE Data Science Data Wrangling and Preprocessing Questions and Answers Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 FREE Data Science Data Wrangling and Preprocessing Questions and Answers flashcards as text
Which technique is most appropriate for handling missing numerical data when the dataset contains significant outliers?
Answer: Median imputation
Median imputation is preferred over mean imputation when outliers are present because the median is robust to extreme values.
What is the primary purpose of one-hot encoding in data preprocessing?
Answer: To convert categorical variables into binary vector representations
One-hot encoding transforms each categorical value into a separate binary column, enabling algorithms that require numerical input to process categorical data.
When performing a log transformation on a feature that contains zero values, what is the standard approach?
Answer: Apply log1p (log of 1 plus the value) instead of a standard log
Log1p adds 1 before taking the logarithm, which handles zero values without producing undefined results while preserving the transformation's variance-stabilizing effect.
In the context of data wrangling, what does the term 'tidy data' refer to?
Answer: Data where each variable is a column, each observation is a row, and each value is a cell
Tidy data follows Hadley Wickham's principles where each variable forms a column, each observation forms a row, and each type of observational unit forms a table.
Which method is most effective for detecting multicollinearity among features during preprocessing?
Answer: Calculating the Variance Inflation Factor (VIF)
VIF quantifies how much the variance of a regression coefficient is inflated due to collinearity with other predictors, with values above 5-10 indicating problematic multicollinearity.
What is the key difference between label encoding and ordinal encoding for categorical variables?
Answer: Ordinal encoding respects a meaningful rank order while label encoding assigns arbitrary integers
Ordinal encoding assigns integers that reflect a meaningful order (e.g., low=1, medium=2, high=3), while label encoding assigns arbitrary integers without implying any rank relationship.