โ† All Data Science Flashcard Decks

FREE Data Science Data Wrangling and Preprocessing Questions and Answers Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 FREE Data Science Data Wrangling and Preprocessing Questions and Answers flashcards as text
  1. Which technique is most appropriate for handling missing numerical data when the dataset contains significant outliers?

    Answer: Median imputation

    Median imputation is preferred over mean imputation when outliers are present because the median is robust to extreme values.

  2. What is the primary purpose of one-hot encoding in data preprocessing?

    Answer: To convert categorical variables into binary vector representations

    One-hot encoding transforms each categorical value into a separate binary column, enabling algorithms that require numerical input to process categorical data.

  3. When performing a log transformation on a feature that contains zero values, what is the standard approach?

    Answer: Apply log1p (log of 1 plus the value) instead of a standard log

    Log1p adds 1 before taking the logarithm, which handles zero values without producing undefined results while preserving the transformation's variance-stabilizing effect.

  4. In the context of data wrangling, what does the term 'tidy data' refer to?

    Answer: Data where each variable is a column, each observation is a row, and each value is a cell

    Tidy data follows Hadley Wickham's principles where each variable forms a column, each observation forms a row, and each type of observational unit forms a table.

  5. Which method is most effective for detecting multicollinearity among features during preprocessing?

    Answer: Calculating the Variance Inflation Factor (VIF)

    VIF quantifies how much the variance of a regression coefficient is inflated due to collinearity with other predictors, with values above 5-10 indicating problematic multicollinearity.

  6. What is the key difference between label encoding and ordinal encoding for categorical variables?

    Answer: Ordinal encoding respects a meaningful rank order while label encoding assigns arbitrary integers

    Ordinal encoding assigns integers that reflect a meaningful order (e.g., low=1, medium=2, high=3), while label encoding assigns arbitrary integers without implying any rank relationship.