← All DA Flashcard Decks

Data Cleaning and Preparation Flashcards

6 cards from real DA practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Data Cleaning and Preparation flashcards as text
  1. What is a left join imputation strategy?

    Answer: Filling NULLs from a reference lookup table via a left join

    Joining a dataset to a reference table on a key and using matched values to fill NULLs is a practical imputation strategy in SQL-based data preparation.

  2. What is the difference between structured and unstructured data?

    Answer: Structured data fits in predefined rows/columns; unstructured data lacks a fixed schema

    Structured data fits neatly into tables (e.g., databases), while unstructured data (emails, images, text) requires processing before it can be analyzed.

  3. What is the purpose of a data dictionary?

    Answer: Documenting the definitions, data types, and relationships of dataset fields

    A data dictionary is a reference document that describes each field in a dataset — its name, type, allowed values, and meaning.

  4. What is the risk of removing rows with missing values (listwise deletion)?

    Answer: It can introduce bias if missing data is not random

    Listwise deletion can bias results if data is not missing completely at random, as systematically removing rows may skew the remaining sample.

  5. What is data reshaping in pandas (Python)?

    Answer: Converting data between wide and long formats using melt or pivot

    Reshaping transforms the layout of a DataFrame, such as using melt() to go from wide to long format or pivot() to do the reverse.

  6. Which of the following is a sign of low-quality data?

    Answer: High percentage of NULL values in critical fields

    A high proportion of NULLs in important columns indicates data quality issues that can undermine analysis accuracy.