← All MS-DS Master of Data science Flashcard Decks

Statistical Inference Concepts Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Statistical Inference Concepts flashcards as text
  1. In a chi-square goodness-of-fit test, the test statistic is large when:

    Answer: Observed frequencies deviate substantially from expected frequencies

    The chi-square statistic Σ(O−E)²/E is large when observed counts differ greatly from what the null hypothesis predicts.

  2. The delta method is used to approximate the variance of g(X̄) when:

    Answer: g is a smooth (differentiable) function and n is large

    The delta method uses a first-order Taylor expansion of g around μ; it requires g to be differentiable and works well for large n via the CLT.

  3. A p-value of 0.001 in a study with n = 10,000 primarily reflects:

    Answer: A very small effect size that is detectable only due to the huge sample

    With n = 10,000, even tiny effects produce very small p-values; the p-value conflates effect size with sample size.

  4. In hypothesis testing, which error is controlled directly by setting the significance level α?

    Answer: Type I error (rejecting a true H₀)

    The significance level α is defined as the maximum acceptable probability of a Type I error (false positive).

  5. The Rao-Blackwell theorem states that if T is a sufficient statistic and U is an unbiased estimator, then E[U|T] is:

    Answer: Unbiased and has variance no greater than that of U

    Conditioning an unbiased estimator on a sufficient statistic yields an unbiased estimator with equal or lower variance — the Rao-Blackwell improvement.

  6. Which of the following correctly describes the relationship between a two-sided 95% confidence interval and a two-sided hypothesis test at α = 0.05?

    Answer: H₀: μ = μ₀ is rejected at α = 0.05 if and only if μ₀ falls outside the 95% CI

    There is a duality: a value μ₀ is in the 95% CI if and only if the test of H₀: μ = μ₀ fails to reject at α = 0.05.

  7. A Bayesian credible interval differs from a frequentist confidence interval in that:

    Answer: A 95% credible interval has a 95% posterior probability of containing the true parameter

    A Bayesian credible interval directly states that P(θ ∈ I | data) = 0.95, a direct probability statement about the parameter given the data.