โ† All Data Science Flashcard Decks

Data Cleaning and Preparation Flashcards

7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Data Cleaning and Preparation flashcards as text
  1. In pandas, df['col'].fillna(method='ffill') does what?

    Answer: Fills missing values with the previous valid value

    Forward fill propagates the last valid observation forward.

  2. Which describes 'data leakage' during preparation?

    Answer: Information from the test set influences training preprocessing

    Leakage occurs when test-set information leaks into training, e.g. fitting a scaler on all data.

  3. To avoid leakage, a StandardScaler should be fit on:

    Answer: The training set only, then applied to test data

    Fit scalers on training data and transform test data with those parameters.

  4. Binning a continuous 'age' column into groups like 'child', 'adult', 'senior' is called:

    Answer: Discretization

    Converting continuous values into discrete bins is discretization.

  5. Which is a sign of a poorly structured ('untidy') dataset?

    Answer: Multiple variables stored in a single column

    Tidy data needs one variable per column; cramming several into one breaks that.

  6. Stripping leading/trailing whitespace and fixing inconsistent capitalization in a text column is part of:

    Answer: Text/string cleaning

    Trimming whitespace and normalizing case are string-cleaning steps.

  7. Median imputation is often preferred over mean imputation when the column:

    Answer: Is skewed or has outliers

    The median is robust to skew and outliers, unlike the mean.