← All MS-DS Master of Data science Flashcard Decks

Statistical Inference Concepts Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Statistical Inference Concepts flashcards as text
  1. The Kolmogorov-Smirnov test is used to:

    Answer: Test whether a sample comes from a specified distribution or compare two samples

    The KS test compares the empirical CDF against a theoretical CDF (one-sample) or between two empirical CDFs (two-sample) to detect distributional differences.

  2. In maximum likelihood estimation, the score function is defined as:

    Answer: The gradient of the log-likelihood with respect to the parameter

    The score is ∂ℓ(θ)/∂θ, the first derivative of the log-likelihood; at the MLE it equals zero.

  3. When using the Wilcoxon signed-rank test instead of a paired t-test, the primary reason is:

    Answer: The Wilcoxon test does not assume normality and is more robust to outliers

    The Wilcoxon signed-rank test is a non-parametric alternative that only assumes symmetry of differences, making it robust when normality is violated.

  4. In the context of regression, what does the term 'heteroscedasticity' refer to?

    Answer: Error variance that changes across levels of the predictors

    Heteroscedasticity means the variance of the regression errors is not constant (σᵢ² varies with X), violating the homoscedasticity assumption of OLS.

  5. Which of the following is the correct formula for the standard error of a proportion p̂ estimated from n observations?

    Answer: √(p̂(1−p̂)/n)

    For a binomial proportion, SE(p̂) = √[p̂(1−p̂)/n], derived from the variance of the Bernoulli distribution divided by n.

  6. A likelihood ratio test statistic −2ln(Λ) follows approximately which distribution under H₀ for large samples?

    Answer: Chi-square distribution with degrees of freedom equal to the number of constrained parameters

    By Wilks' theorem, −2ln(Λ) is asymptotically χ²(k) where k is the number of parameters constrained under H₀.

  7. Which statement best describes the concept of a conjugate prior in Bayesian inference?

    Answer: A prior from the same distributional family as the posterior, simplifying analytical computation

    A conjugate prior yields a posterior in the same parametric family, making closed-form Bayesian updating possible (e.g., Beta prior with Binomial likelihood gives a Beta posterior).