Data Cleaning and Preparation Flashcards
7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Cleaning and Preparation flashcards as text
In pandas, df['col'].fillna(method='ffill') does what?
Answer: Fills missing values with the previous valid value
Forward fill propagates the last valid observation forward.
Which describes 'data leakage' during preparation?
Answer: Information from the test set influences training preprocessing
Leakage occurs when test-set information leaks into training, e.g. fitting a scaler on all data.
To avoid leakage, a StandardScaler should be fit on:
Answer: The training set only, then applied to test data
Fit scalers on training data and transform test data with those parameters.
Binning a continuous 'age' column into groups like 'child', 'adult', 'senior' is called:
Answer: Discretization
Converting continuous values into discrete bins is discretization.
Which is a sign of a poorly structured ('untidy') dataset?
Answer: Multiple variables stored in a single column
Tidy data needs one variable per column; cramming several into one breaks that.
Stripping leading/trailing whitespace and fixing inconsistent capitalization in a text column is part of:
Answer: Text/string cleaning
Trimming whitespace and normalizing case are string-cleaning steps.
Median imputation is often preferred over mean imputation when the column:
Answer: Is skewed or has outliers
The median is robust to skew and outliers, unlike the mean.