Data Science Statistical Concepts and Inference Questions and Answers Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Data Science Statistical Concepts and Inference Questions and Answers flashcards as text
A data scientist obtains a 95% confidence interval of (12.3, 18.7) for the population mean. Which interpretation is correct?
Answer: If we repeated sampling many times, 95% of constructed intervals would contain the true mean
A 95% confidence interval means that 95% of intervals constructed from repeated samples would capture the true population parameter.
Which of the following scenarios would most likely violate the assumption of independence required for a two-sample t-test?
Answer: Comparing blood pressure readings before and after treatment on the same patients
Measuring the same patients before and after creates paired observations that are not independent, requiring a paired t-test instead.
In hypothesis testing, what does a Type II error represent?
Answer: Failing to reject a false null hypothesis
A Type II error occurs when we fail to reject the null hypothesis even though it is actually false, meaning we miss a real effect.
A researcher increases the sample size from 50 to 200 while keeping everything else constant. What is the most direct effect on the standard error of the mean?
Answer: It is cut in half
Standard error equals sigma divided by the square root of n, so quadrupling n from 50 to 200 halves the standard error.
When performing multiple comparisons across 20 groups using individual t-tests at alpha = 0.05, what problem arises?
Answer: The familywise error rate inflates well beyond 0.05
Running many pairwise tests without correction inflates the overall probability of at least one false positive far above the nominal alpha level.
Which measure of central tendency is most robust to outliers in a skewed dataset?
Answer: Median
The median is the middle value when data is sorted and is unaffected by extreme values, making it robust to outliers in skewed distributions.