โ† All Data Science Flashcard Decks

Data Wrangling and Preprocessing Flashcards

7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Data Wrangling and Preprocessing flashcards as text
  1. When concatenating training feature engineering, fitting an encoder on the full dataset before train/test split causes what problem?

    Answer: Data leakage

    Fitting any preprocessing on the combined data lets test information influence training, which is data leakage.

  2. A target variable has 95% class A and 5% class B. Which preprocessing technique addresses this imbalance?

    Answer: Resampling such as SMOTE or undersampling

    Oversampling the minority class (e.g., SMOTE) or undersampling the majority class rebalances skewed class distributions.

  3. Why should imputation of missing values inside a cross-validation loop be done within each fold rather than once beforehand?

    Answer: To avoid leaking information across folds

    Imputing before splitting lets statistics from validation folds leak into training, inflating performance estimates.

  4. Two highly correlated numeric features (r=0.98) are kept in a linear model. What issue does this raise?

    Answer: Multicollinearity

    Highly correlated predictors cause multicollinearity, making coefficient estimates unstable and hard to interpret.

  5. A free-text 'notes' column needs to become model features. Which preprocessing step is most appropriate first?

    Answer: Text vectorization (e.g., TF-IDF or tokenization)

    Raw text must be tokenized and vectorized (e.g., TF-IDF) into numeric features before most models can use it.

  6. When using scikit-learn, bundling scaling, encoding, and the model into a single object that fits in order prevents leakage and simplifies deployment. What is it called?

    Answer: A Pipeline

    A Pipeline chains preprocessing steps and the estimator so transforms are fit only on training data within each fold.

  7. You scale numeric columns but leave one-hot encoded binary columns unscaled. Why is this generally acceptable?

    Answer: Binary 0/1 columns are already on a comparable bounded scale

    One-hot columns are already bounded to 0 and 1, so they are on a scale comparable to standardized numeric features.