Data Science Data Cleaning and Preparation Questions and Answers Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Data Science Data Cleaning and Preparation Questions and Answers flashcards as text
What is the main advantage of using K-Nearest Neighbors imputation over simple mean imputation for missing data?
Answer: It considers relationships between features to produce context-aware estimates
KNN imputation leverages feature correlations by finding similar records, producing more realistic imputed values than a single global mean.
When merging two datasets with different date formats (MM/DD/YYYY and YYYY-MM-DD), what should be done first?
Answer: Parse both columns into a standardized datetime format before merging
Standardizing date formats before merging prevents mismatches, join failures, and incorrect date comparisons.
Which technique helps detect duplicate records that have slight variations in text fields such as customer names?
Answer: Fuzzy string matching using algorithms like Levenshtein distance
Fuzzy matching algorithms measure string similarity and can identify near-duplicates like 'Jon Smith' and 'John Smith' that exact matching would miss.
What problem does multicollinearity cause during data preparation for regression models?
Answer: Unstable coefficient estimates that change dramatically with small data variations
Multicollinearity inflates the variance of coefficient estimates, making them unreliable and sensitive to minor changes in the data.
What is the purpose of applying a Box-Cox transformation during data preparation?
Answer: To make a skewed distribution more closely approximate a normal distribution
Box-Cox transformation applies a power transformation that reduces skewness and helps data better satisfy normality assumptions required by many statistical methods.
When splitting data into training and test sets, why should stratified sampling be used for imbalanced classification problems?
Answer: To ensure both sets maintain the same class distribution as the original data
Stratified sampling preserves the original class proportions in both splits, preventing situations where the minority class is underrepresented or absent in one set.