Data Cleaning and Preparation Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Data Cleaning and Preparation flashcards as text
The interquartile range (IQR) method is commonly used to detect:
Answer: Outliers
Values outside 1.5×IQR beyond the quartiles are flagged as outliers.
Min-max scaling transforms a feature to which range by default?
Answer: [0, 1]
Min-max scaling rescales values to the 0 to 1 range.
Z-score standardization produces data with:
Answer: Mean of 0 and standard deviation of 1
Standardization centers data to mean 0 and unit standard deviation.
Which encoding creates a separate binary column for each category value?
Answer: One-hot encoding
One-hot encoding generates one binary indicator column per category.
Why scale features before training a k-nearest-neighbors model?
Answer: Distance calculations are sensitive to feature magnitude
KNN uses distances, so unscaled large-magnitude features dominate.
Applying a log transform to a right-skewed feature primarily aims to:
Answer: Reduce skew and compress large values
Log transforms compress large values and reduce right skew.
Removing a data point with age = 250 from a human dataset is best described as:
Answer: Handling an invalid/implausible value
An impossible age is an invalid value to correct or remove.