← All CBDA Flashcard Decks

Statistical Analysis & Methods Flashcards

6 cards from real CBDA practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Statistical Analysis & Methods flashcards as text
  1. What is the primary purpose of a confidence interval in business data analysis?

    Answer: To provide a range within which the true population parameter likely falls

    A confidence interval estimates the range that contains the true population parameter with a specified level of certainty (e.g., 95%).

  2. In regression analysis, what do 'residuals' represent?

    Answer: The difference between actual observed values and the model's predicted values

    Residuals are the errors — the gap between what the model predicted and what was actually observed.

  3. A CBDA analyst applies z-score normalization to a dataset. What is the MAIN reason for doing this?

    Answer: To scale variables to a common mean of 0 and standard deviation of 1 for comparability

    Z-score normalization standardizes variables so they can be meaningfully compared or used together in models regardless of original scale.

  4. Which of the following describes a statistically significant result that is NOT practically significant?

    Answer: A very small real-world effect detected as significant because of a very large sample size

    With very large samples, even trivially small effects become statistically significant, but may have no meaningful business impact.

  5. What is 'overfitting' in the context of predictive modeling for business data analysts?

    Answer: Building a model that performs well on training data but poorly on new, unseen data

    An overfit model has learned the noise in the training data rather than the underlying pattern, causing poor generalization.

  6. A CBDA analyst uses cross-validation when evaluating a predictive model. What is the main benefit of this technique?

    Answer: It provides a more reliable estimate of model performance by testing on multiple held-out data subsets

    Cross-validation rotates the hold-out set across multiple folds, producing a more robust and unbiased estimate of true model performance.