Exploratory Data Analysis Techniques Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Exploratory Data Analysis Techniques flashcards as text
Which plot is most suitable for visualizing changes in a continuous variable over time?
Answer: Line chart
A line chart connects sequential data points in chronological order, making trends, cycles, and anomalies in time-series data easy to see.
What is the purpose of a pivot table in EDA?
Answer: To summarize data by aggregating values across two or more categorical dimensions
A pivot table reorganizes data into a matrix where rows and columns represent categorical variables and cells contain aggregated statistics like counts or means.
A dataset has a bimodal distribution. What does this likely indicate?
Answer: The dataset may contain two distinct subpopulations or groups
Two peaks (modes) in a distribution often suggest the presence of two distinct subgroups with different characteristic values mixed into a single dataset.
Which of the following is a key step in univariate EDA for a continuous variable?
Answer: Examining the distribution shape, central tendency, spread, and outliers
Univariate EDA for a continuous variable involves summarizing and visualizing its distribution shape, mean/median, standard deviation, and extreme values.
What does Spearman's rank correlation measure that Pearson's correlation does not?
Answer: Monotonic relationships, including nonlinear ones, between two variables
Spearman's correlation computes Pearson's r on the ranks of the data, capturing any monotonic relationship (not just linear) and being robust to outliers.
In EDA, what does 'cardinality' of a categorical variable refer to?
Answer: The number of unique distinct values the variable can take
Cardinality is the count of distinct categories in a categorical variable; high-cardinality variables (e.g., user IDs) may need special encoding strategies.
Which technique helps identify the most important variables early in EDA without building a full model?
Answer: Computing correlation with the target variable or using mutual information scores
Correlations and mutual information scores quantify how much each feature relates to the target, providing a fast model-free signal for feature relevance.