โ† All Data Science Flashcard Decks

Data Wrangling and Preprocessing Flashcards

7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Data Wrangling and Preprocessing flashcards as text
  1. Two rows have identical values across all columns. What wrangling step removes them?

    Answer: Deduplication (drop_duplicates)

    Dropping duplicates removes repeated rows that would otherwise bias counts and model training.

  2. A 'state' column contains 'CA', 'Ca', 'calif.', and 'California' for the same state. This is an example of what problem?

    Answer: Inconsistent categorical values needing standardization

    The same category written multiple ways must be standardized to a single canonical value before analysis.

  3. Using the IQR method, an outlier is typically a value beyond which boundary?

    Answer: Below Q1 - 1.5*IQR or above Q3 + 1.5*IQR

    The Tukey IQR rule flags points more than 1.5 times the interquartile range beyond the first or third quartile.

  4. You have monthly sales spread across 12 separate columns (Jan...Dec). Reshaping into a single 'month' column and 'sales' column is called what?

    Answer: Melting / unpivoting to long format

    Melting (unpivoting) converts wide-format columns into long-format key-value rows.

  5. Converting a continuous 'age' feature into groups like 'child', 'adult', 'senior' is known as what?

    Answer: Binning (discretization)

    Binning discretizes a continuous variable into a finite set of intervals or categories.

  6. Why is it risky to drop an entire column that has 60% missing values without investigation?

    Answer: The missingness itself may carry predictive signal

    Missingness can be informative (e.g., not-applicable cases), so dropping blindly can discard a useful signal.

  7. Which pandas operation would you use to apply a custom cleaning function to every element of a Series?

    Answer: Series.apply() / map()

    apply() and map() run a function element-wise over a Series for transformation or cleaning.