← All CAS Flashcard Decks

Data Analysis and Interpretation Flashcards

7 cards from real CAS practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Data Analysis and Interpretation flashcards as text
  1. An actuary applies a one-way analysis of an insurance rating factor and finds that relativities are not monotone. Before making adjustments, which of the following is the most important consideration?

    Answer: Whether the non-monotone pattern reflects genuine risk or is due to sampling variability or sparse data

    Non-monotone relativities often result from thin data in certain cells; credibility weighting or smoothing may be appropriate before forcing monotonicity.

  2. The coefficient of variation (CV) of a loss distribution is defined as:

    Answer: Standard deviation divided by the mean

    The CV = σ/μ measures relative dispersion; it is dimensionless and allows comparison of variability across distributions with different scales.

  3. When using gradient boosting machines (GBM) for loss cost modeling, the learning rate (shrinkage) hyperparameter primarily controls:

    Answer: The contribution of each successive tree to the overall prediction

    The learning rate scales each tree's contribution, trading off slower learning (smaller rate, more trees needed) for better generalization.

  4. A p-p plot (probability-probability plot) is used to compare an empirical distribution to a theoretical one. Unlike a Q-Q plot, the p-p plot compares:

    Answer: Empirical CDFs evaluated at each data point against theoretical CDF values

    A p-p plot maps the empirical CDF value at each observation against the theoretical CDF value at the same point; deviations from the diagonal indicate distributional misfit.

  5. In a stochastic claims reserving model using bootstrapping of Pearson residuals, the main purpose is to:

    Answer: Generate a distribution of reserve estimates to quantify reserve uncertainty

    Bootstrapping the residuals of a chain-ladder model produces many simulated triangles, yielding a distribution of possible ultimate losses that quantifies reserve risk.

  6. The Cramér-von Mises statistic is used in goodness-of-fit testing. Compared to the Kolmogorov-Smirnov statistic, it:

    Answer: Measures the integrated squared difference between empirical and theoretical CDFs

    The Cramér-von Mises statistic integrates the squared deviation between the empirical and theoretical CDFs over all observed values, giving global rather than maximum deviation.

  7. An insurer builds a classification tree to predict whether a policy will generate a large loss. The tree is grown to full depth on training data (zero misclassification error). The primary problem with deploying this model is:

    Answer: The model is overfit and will likely perform poorly on new, unseen data

    A fully grown tree memorizes training data noise (overfitting), resulting in high variance and degraded predictive performance on hold-out or future data.