← All MS-DS Master of Data science Flashcard Decks

Exploratory Data Analysis Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Exploratory Data Analysis flashcards as text
  1. Which EDA plot is MOST effective for visualizing the joint distribution of two continuous variables, including the density in dense regions?

    Answer: Hex bin plot

    Hex bin plots aggregate points into hexagonal bins colored by count, avoiding overplotting in dense regions where scatter plots become unreadable.

  2. What does it mean when the mean absolute deviation (MAD) of a feature is very close to zero?

    Answer: Almost all values are clustered near the mean with little spread

    A near-zero MAD indicates that observations deviate very little from the mean on average, implying the feature has very low variability.

  3. During EDA, you find that a timestamp column contains dates ranging from 1900 to 2025, but the data should only span 2010–2025. What EDA action should you take first?

    Answer: Visualize the temporal distribution and investigate the anomalous early dates before deciding on treatment

    Investigating the anomalous values first reveals whether they are data entry errors, system defaults, or legitimate historical records before any transformation.

  4. Which measure is used to quantify the 'peakedness' or tail weight of a distribution relative to a normal distribution?

    Answer: Kurtosis

    Kurtosis measures the weight of a distribution's tails; excess kurtosis > 0 (leptokurtic) indicates heavier tails than normal, while < 0 (platykurtic) indicates lighter tails.

  5. A parallel coordinates plot is used in EDA primarily to:

    Answer: Visualize multivariate data and identify patterns or clusters across many dimensions simultaneously

    Parallel coordinates draw each observation as a polyline across multiple parallel axes, making it possible to spot clusters, outliers, and correlations in high-dimensional data.

  6. When should you prefer a log scale on a plot axis during EDA?

    Answer: When the data spans several orders of magnitude

    Log scales compress large ranges so that patterns across orders of magnitude (e.g., 1 to 1,000,000) become visible without small values being crushed near zero.

  7. What is the purpose of computing a 'data profile' (e.g., using pandas-profiling or ydata-profiling) as part of EDA?

    Answer: To generate a comprehensive automated summary of each feature's type, missing values, distributions, and correlations

    Profiling tools auto-generate feature-level statistics, histograms, missing value counts, and correlation matrices, giving analysts a fast holistic overview before deep-dive EDA.