โ† All MS-DS Master of Data science Flashcard Decks

MS-DS Master of Data science Data Wrangling and Preprocessing Questions and Answers Flashcards

6 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 MS-DS Master of Data science Data Wrangling and Preprocessing Questions and Answers flashcards as text
  1. When merging two datasets with a many-to-many relationship on the join key, what is the most critical preprocessing step?

    Answer: Deduplicating or reshaping one dataset to establish a clear one-to-many relationship

    Many-to-many joins produce a Cartesian product of matching rows, so deduplicating or reshaping ensures the merged result has the intended granularity.

  2. Which encoding strategy is most appropriate for a categorical variable with a natural ordering such as education level?

    Answer: Ordinal encoding that preserves the rank order

    Ordinal encoding assigns integers that reflect the inherent order of categories, preserving the ranking information that other encoding methods would discard.

  3. What problem does target encoding introduce that requires careful mitigation through cross-validation or smoothing?

    Answer: Data leakage from the target variable into the features

    Target encoding replaces categories with the mean of the target, which can leak target information into features and cause overfitting without proper regularization.

  4. During text data preprocessing for a data science pipeline, what is the correct order of operations?

    Answer: Lowercasing, tokenization, stop word removal, stemming or lemmatization

    The standard NLP preprocessing pipeline normalizes case first, splits text into tokens, removes uninformative stop words, and then reduces words to their base forms.

  5. What is the primary risk of applying feature scaling before splitting data into training and test sets?

    Answer: Information from the test set leaks into the training set through shared scaling parameters

    Fitting the scaler on the entire dataset before splitting allows test set statistics to influence training data transformation, constituting data leakage.

  6. Which resampling technique for handling class imbalance generates synthetic minority samples by interpolating between existing minority observations?

    Answer: SMOTE (Synthetic Minority Over-sampling Technique)

    SMOTE creates new synthetic examples by interpolating between nearest neighbors in the minority class, increasing representation without simple duplication.