Data Analysis & Statistics Flashcards
6 cards from real Bluebook SAT Test practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Data Analysis & Statistics flashcards as text
A researcher collects data on weekly study hours and exam scores for 200 students. The correlation coefficient is r = 0.72. A student argues that studying more causes higher scores because the correlation is strong. Which statistical concept most directly refutes this causal claim?
Answer: Confounding variables — a third factor such as prior academic ability may independently drive both study hours and scores
Correlation does not imply causation. A confounding variable (e.g., baseline academic ability) could independently increase both study hours and exam performance, producing a strong correlation without a direct causal link. Regression toward the mean and r² misunderstand the logical issue, and n = 200 is sufficient to estimate correlation reliably.
Two data sets each have a mean of 50. Set A has a standard deviation of 2; Set B has a standard deviation of 10. A value of 54 is drawn from one of the sets. Which statement is most accurate?
Answer: The value is more unusual in Set A because 54 is 2 standard deviations above Set A's mean but only 0.4 standard deviations above Set B's mean
Unusualness is measured by z-score: z = (x − μ) / σ. For Set A: z = (54 − 50) / 2 = 2.0. For Set B: z = (54 − 50) / 10 = 0.4. A z-score of 2.0 is much further into the tail than 0.4, so the value is more unusual relative to Set A. The absolute distance from the mean is the same, but standard deviations provide the proper scale.
A scatterplot of x and y shows a clear curved (quadratic) pattern with no linear trend. A student computes the Pearson correlation coefficient and gets r ≈ 0.05. She concludes there is no relationship between x and y. What is wrong with her conclusion?
Answer: Pearson's r only measures the strength of a linear relationship; a near-zero r is consistent with a strong nonlinear association
Pearson's r quantifies the strength and direction of a LINEAR relationship only. A symmetric parabolic (quadratic) relationship can produce r ≈ 0 because positive and negative deviations cancel, even when the association is very strong. The student should use a scatterplot and nonlinear models, not Pearson's r alone, to characterize the relationship.
A polling organization surveys 1,000 randomly selected voters and finds that 54% support a ballot measure, with a reported margin of error of ±3 percentage points at the 95% confidence level. The organization doubles its sample size to 4,000. What happens to the margin of error?
Answer: It is reduced to approximately ±2.1 percentage points
The margin of error is proportional to 1/√n. When sample size quadruples (×4), the margin of error is divided by √4 = 2, giving ±3 / 2 = ±1.5 percentage points — wait, that matches A. Let's recheck: original n = 1,000 → ME ≈ 3%. New n = 4,000. ME scales by √(1000/4000) = √(1/4) = 1/2. New ME ≈ 1.5%. The correct answer is actually A (≈ ±1.5 pp). Answer B (≈ ±2.1 pp) would result from doubling, not quadrupling, the sample (√2 ≈ 1.41, 3/1.41 ≈ 2.1). The correct choice here is A.
A box plot shows the following five-number summary for a dataset: Min = 10, Q1 = 30, Median = 50, Q3 = 70, Max = 130. Using the standard 1.5 × IQR rule, which values, if any, are outliers?
Answer: 130 is an outlier; 10 is not
IQR = Q3 − Q1 = 70 − 30 = 40. Lower fence = Q1 − 1.5 × IQR = 30 − 60 = −30. Upper fence = Q3 + 1.5 × IQR = 70 + 60 = 130. The value 130 sits exactly on the upper fence. By the strict rule, a point must be BEYOND the fence (> 130 or −30 (not an outlier); 130 = 130 (boundary case). The SAT Bluebook convention marks values that EXCEED the fences. Therefore neither is an outlier, and the correct answer is C.
A study reports that students who ate breakfast daily scored an average of 8 points higher on a standardized test than those who skipped breakfast. The researchers randomly assigned 500 students to eat breakfast and 500 to skip it for one month before testing. What is the most valid conclusion?
Answer: Eating breakfast likely causes higher test scores because the study design controls for confounding variables through random assignment
This study is a randomized controlled experiment, not an observational study. Random assignment of participants to treatment (breakfast) and control (no breakfast) groups controls for confounding variables, allowing researchers to draw causal conclusions. The 8-point difference can be attributed to the breakfast intervention itself. Answer D is incorrect because participants were assigned, not self-selected.