Data Analytics Flashcards
7 cards from real CAIC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Data Analytics flashcards as text
An AI consultant is designing a feature engineering pipeline. Which transformation converts a high-cardinality categorical variable (e.g., ZIP codes) into numeric inputs without creating thousands of sparse columns?
Answer: Target encoding
Target encoding replaces each category with the mean of the target variable for that category, capturing signal without the dimensionality explosion of one-hot encoding.
What is the primary risk of using test data to tune model hyperparameters?
Answer: Data leakage causing overly optimistic performance estimates
Using test data for tuning leaks future information into the modeling process, inflating performance metrics and making the model appear better than it will perform on truly unseen data.
Which metric is most informative when evaluating a ranking model (e.g., search results or recommendations)?
Answer: Normalized Discounted Cumulative Gain (NDCG)
NDCG measures ranking quality by rewarding models that place highly relevant results near the top, discounting relevance gains for lower-ranked positions.
A stakeholder requests that all model predictions be explainable to regulators. Which framework provides model-agnostic local explanations for individual predictions?
Answer: SHAP (SHapley Additive exPlanations)
SHAP assigns each feature a contribution value for a specific prediction using game-theory-based Shapley values, working with any model type.
When performing exploratory data analysis (EDA), what is the main purpose of computing pairwise correlation between numeric features?
Answer: To detect multicollinearity and relationships that could affect model training
Pairwise correlation in EDA reveals multicollinearity (highly correlated predictors) and feature-target relationships that inform feature selection decisions.
A company stores user event logs in a data lake but queries are very slow. An AI consultant recommends partitioning the data. What is the primary benefit?
Answer: It reduces query scan volume by limiting reads to relevant data partitions
Partitioning (e.g., by date or region) allows query engines to skip irrelevant partitions entirely, dramatically reducing I/O and speeding up analytical queries.
Which practice is essential before deploying an AI model to production to ensure it behaves consistently across subgroups defined by race, gender, or age?
Answer: Conducting a bias and fairness audit using disaggregated performance metrics
A bias and fairness audit evaluates model performance separately for each protected subgroup to identify disparate impact before deployment.