Data Science Data Cleaning and Preparation Questions and Answers Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Data Science Data Cleaning and Preparation Questions and Answers flashcards as text
Which technique is most appropriate for handling missing values in a time series dataset?
Answer: Forward fill or interpolation
Forward fill and interpolation preserve temporal continuity, making them ideal for time series data where adjacent values are related.
What is the primary risk of removing outliers without domain knowledge?
Answer: Losing valid but extreme observations that represent real phenomena
Outliers may represent genuine rare events, and removing them without understanding the domain can eliminate important signal from the data.
When performing one-hot encoding on a categorical variable with 50 unique values, what problem is most likely to arise?
Answer: The curse of dimensionality from creating too many sparse features
One-hot encoding a high-cardinality feature creates 50 new binary columns, dramatically increasing dimensionality and sparsity.
What does the term 'data leakage' refer to in the context of data preparation?
Answer: Information from outside the training set improperly influencing model building
Data leakage occurs when information that would not be available at prediction time is used during training, leading to overly optimistic performance estimates.
Which method is best suited for normalizing a feature that contains significant outliers?
Answer: Robust scaling using median and interquartile range
Robust scaling uses the median and IQR, which are resistant to outliers, unlike min-max or z-score methods that are heavily influenced by extreme values.
What is the correct order of operations when preparing a dataset with missing values and features requiring scaling?
Answer: Impute missing values first, then apply feature scaling
Imputation must occur before scaling because most scaling methods cannot handle missing values and would either error or produce misleading results.