Data Analytics Flashcards
7 cards from real CAIC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Data Analytics flashcards as text
Which statistical technique is most appropriate for identifying hidden groupings in unlabeled customer data?
Answer: K-means clustering
K-means clustering is an unsupervised technique that partitions data into k groups based on feature similarity without requiring labels.
An AI consultant is asked to evaluate model fairness. Which metric detects whether a model performs significantly worse for a protected demographic group?
Answer: Disparate impact ratio
Disparate impact ratio compares outcome rates across demographic groups and flags values below 0.8 as potentially discriminatory.
What does a high variance but low bias in a predictive model typically indicate?
Answer: The model is overfitting
High variance and low bias indicates the model has learned noise in the training data and fails to generalize—a classic overfitting pattern.
A data pipeline silently drops records with NULL values in a key field. What type of data quality issue does this represent?
Answer: Silent data loss
Silent data loss occurs when records are discarded without alerts or logging, making the issue invisible to downstream consumers.
Which visualization type is best suited for showing the distribution and spread of a continuous numeric variable?
Answer: Box plot
A box plot displays the median, quartiles, and outliers of a continuous variable, making distribution and spread immediately visible.
When building an AI recommendation system, which data analytics step ensures the training data reflects current user behavior rather than outdated patterns?
Answer: Data freshness validation
Data freshness validation checks that training data timestamps are recent enough to reflect current behavioral patterns before model training.
A model's precision is 0.90 but recall is 0.40 on a fraud detection task. What is the best interpretation?
Answer: The model misses most fraud cases but rarely flags legitimate transactions as fraud
Low recall (0.40) means 60% of actual fraud is missed, while high precision (0.90) means flagged cases are usually correct—the model is too conservative.