โ† All MS-DS Master of Data science Flashcard Decks

Master of Data science Data Wrangling and Preprocessing 1 Flashcards

6 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Master of Data science Data Wrangling and Preprocessing 1 flashcards as text
  1. Which pandas function is used to detect missing values in a DataFrame, returning a boolean mask?

    Answer: df.isnull()

    df.isnull() returns a DataFrame of the same shape with True where values are NaN and False elsewhere, making it the standard method for detecting missing data.

  2. What is the main difference between StandardScaler and MinMaxScaler in scikit-learn?

    Answer: StandardScaler transforms features to zero mean and unit variance; MinMaxScaler scales features to a fixed range like [0,1]

    StandardScaler centers data by subtracting the mean and dividing by standard deviation, producing zero mean and unit variance. MinMaxScaler maps values to a bounded range (default [0,1]) by subtracting the minimum and dividing by the range.

  3. In pandas, what is the effect of setting the parameter how='outer' in pd.merge()?

    Answer: It returns all rows from both DataFrames, filling NaN where there is no match

    An outer join returns the union of keys from both DataFrames. Where a key exists in one DataFrame but not the other, the missing columns are filled with NaN.

  4. Which technique is best suited for imputing missing values in a numerical column when the data is believed to be missing at random and the column has a moderate skew?

    Answer: Median imputation

    Median imputation is preferred over mean imputation when data is skewed because the median is robust to outliers and better represents the central tendency of a skewed distribution.

  5. What does the pandas groupby() method return before an aggregation function is applied?

    Answer: A DataFrameGroupBy object

    df.groupby() returns a DataFrameGroupBy object, which is a lazy grouping structure. The actual computation only occurs when an aggregation function such as .sum(), .mean(), or .agg() is called on it.

  6. When applying label encoding to an ordinal categorical feature with categories ['Low', 'Medium', 'High'], which concern is most important to address?

    Answer: The assigned integer values must preserve the natural order of the categories

    For ordinal features, the integer labels assigned by encoding must reflect the correct ordering (e.g., Low=0, Medium=1, High=2). Using arbitrary or reversed numeric assignments would mislead models that interpret numeric magnitude as meaningful.