โ† All Data Science Flashcard Decks

Data Cleaning and Preparation Flashcards

7 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Data Cleaning and Preparation flashcards as text
  1. In pandas, what does df.dropna(thresh=3) do?

    Answer: Drops rows that have fewer than 3 non-null values

    thresh keeps only rows with at least 3 non-null values, dropping those below the threshold.

  2. Which technique replaces missing numeric values with the column average?

    Answer: Mean imputation

    Mean imputation fills missing values using the column's average.

  3. A 'duplicate record' in a dataset is best defined as a row that:

    Answer: Has identical values across the columns used to identify uniqueness

    Duplicates are rows matching on the key columns that define a unique record.

  4. Which pandas method removes duplicate rows?

    Answer: df.drop_duplicates()

    drop_duplicates() removes repeated rows from a DataFrame.

  5. Converting a column stored as text '2024-01-15' into a datetime type is an example of:

    Answer: Type casting / parsing

    Changing a string into a datetime is type casting or parsing.

  6. What is the main risk of dropping all rows containing any missing value?

    Answer: Significant loss of usable data and potential bias

    Listwise deletion can discard large amounts of data and introduce bias.

  7. Standardizing inconsistent category labels like 'USA', 'U.S.A.', and 'United States' is called:

    Answer: Data standardization / canonicalization

    Mapping variant labels to one canonical form is standardization.