โ† All Data Science Flashcard Decks

Data Cleaning and Preparation Flashcards

7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Data Cleaning and Preparation flashcards as text
  1. Combining two DataFrames on a shared key column in pandas is done with:

    Answer: pd.merge()

    pd.merge joins DataFrames on common key columns.

  2. An 'inner join' between two tables returns:

    Answer: Only rows with matching keys in both tables

    An inner join keeps only keys present in both tables.

  3. MICE (Multiple Imputation by Chained Equations) is a method for:

    Answer: Imputing missing values using other features

    MICE models each feature with missing data from the others to impute values.

  4. Reshaping data from wide to long format in pandas typically uses:

    Answer: pd.melt()

    pd.melt unpivots wide columns into long key-value rows.

  5. When categorical labels have a natural order (low, medium, high), the best encoding is:

    Answer: Ordinal encoding

    Ordinal encoding preserves the inherent rank order of categories.

  6. A 'data dictionary' in a preparation workflow primarily documents:

    Answer: Each field's meaning, type, and allowed values

    A data dictionary defines column meanings, types, and valid values.

  7. SMOTE is applied during preparation mainly to address:

    Answer: Class imbalance by synthesizing minority-class samples

    SMOTE oversamples the minority class by creating synthetic examples.