Data Analytics Flashcards
7 cards from real CAIC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Data Analytics flashcards as text
Which technique is used to estimate model performance on unseen data when the full dataset is too small to hold out a separate test set?
Answer: k-fold cross-validation
k-fold cross-validation splits data into k subsets, trains on k-1 folds and validates on the remaining fold, rotating through all folds for a robust performance estimate.
A business analyst notices that sales data always spikes every December regardless of the year. This pattern is best described as:
Answer: Seasonality
Seasonality refers to periodic, predictable fluctuations that repeat at fixed intervals such as monthly or yearly cycles.
An AI model trained on pre-pandemic data is deployed post-pandemic and performs poorly. This failure is best explained by:
Answer: Concept drift due to changed real-world relationships
Concept drift occurs when the statistical relationship between inputs and outputs changes over time, making previously valid models obsolete.
Which SQL window function would an analyst use to calculate a running total of sales ordered by date without collapsing rows?
Answer: SUM() OVER (ORDER BY date)
SUM() OVER (ORDER BY date) is a window function that computes a cumulative sum per row while preserving all rows in the result set.
When should an AI consultant recommend using a data warehouse over a transactional database for analytics?
Answer: When complex historical queries and reporting across large datasets are needed
Data warehouses are optimized for analytical workloads—columnar storage and query optimization make them far more efficient for aggregations over large historical data than OLTP systems.
A model achieves 99% accuracy on a dataset where 99% of records belong to class A. What does this reveal?
Answer: Accuracy is misleading; the model likely predicts class A for everything
In highly imbalanced datasets, a trivial classifier that always predicts the majority class can achieve deceptively high accuracy while completely failing the minority class.
Which approach best addresses the 'cold start' problem in a collaborative filtering recommendation system?
Answer: Using content-based features for new users or items with no interaction history
Content-based filtering leverages item or user attributes to generate recommendations when no interaction data exists, solving the cold start problem.