← All CAS Flashcard Decks

Data Analysis and Interpretation Flashcards

7 cards from real CAS practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Data Analysis and Interpretation flashcards as text
  1. When fitting a generalized linear model (GLM) to insurance loss data, which link function is most commonly used for modeling claim frequency?

    Answer: Log link

    The log link is standard for Poisson-distributed claim frequency models because it ensures predicted frequencies remain positive.

  2. A claims analyst observes that residuals from a regression model exhibit a funnel shape when plotted against fitted values. This pattern most likely indicates:

    Answer: Heteroscedasticity in the error terms

    A funnel-shaped residual plot is the classic diagnostic for heteroscedasticity, where error variance is not constant across fitted values.

  3. In credibility theory, the Bühlmann credibility factor Z is defined as n/(n+k). As the number of observations n increases toward infinity, Z approaches:

    Answer: 1

    As n → ∞, the credibility factor Z = n/(n+k) approaches 1, meaning full weight is given to observed experience.

  4. Which of the following best describes the use of a Q-Q plot in actuarial data analysis?

    Answer: Assessing whether data follow a specified theoretical distribution

    A Q-Q plot compares empirical quantiles of data against theoretical quantiles of a reference distribution to assess distributional fit.

  5. An insurer uses principal component analysis (PCA) on 20 rating variables. The first three principal components explain 85% of total variance. The primary benefit of using these three components instead of all 20 is:

    Answer: Reducing dimensionality while retaining most variance

    PCA's main benefit is dimensionality reduction — capturing most variance with far fewer uncorrelated components, which simplifies modeling.

  6. A dataset of 500 claims has a sample mean of $12,000 and sample standard deviation of $8,000. The 95% confidence interval for the population mean is constructed using the t-distribution rather than the normal because:

    Answer: The population standard deviation is unknown

    The t-distribution is used when the population standard deviation is unknown and must be estimated from sample data.

  7. In a loss development triangle, the volume-weighted average link ratio for a development period is calculated by:

    Answer: Dividing total cumulative losses at the later age by total cumulative losses at the earlier age

    The volume-weighted (chain-ladder) link ratio divides the sum of later-age cumulative losses by the sum of earlier-age cumulative losses, giving more weight to larger values.