Data Analysis & Statistics Flashcards
6 cards from real Bluebook SAT Test practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Data Analysis & Statistics flashcards as text
The line of best fit for a scatterplot of hours studied (x) versus exam score (y) is ŷ = 4.2x + 55. A student who studied 8 hours received a score of 87. What is the residual for this data point, and what does it indicate about the model's prediction?
Answer: −1.6; the student performed below the model's prediction
The predicted score is ŷ = 4.2(8) + 55 = 33.6 + 55 = 88.6. The residual = actual − predicted = 87 − 88.6 = −1.6. A negative residual means the actual score fell below the model's prediction — the line overestimates this student's score, not underestimates it. Choice C reverses the interpretation of a negative residual.
Dataset X has a mean of 50 and a standard deviation of 2. Dataset Y has a mean of 200 and a standard deviation of 8. Which statement correctly compares the relative variability of the two datasets?
Answer: Both datasets have equal relative variability
Relative variability is measured by the coefficient of variation (CV = standard deviation ÷ mean). For Dataset X: CV = 2/50 = 0.04 (4%). For Dataset Y: CV = 8/200 = 0.04 (4%). Despite Dataset Y having a larger absolute standard deviation, both datasets spread by the same proportion relative to their respective means. Comparing raw standard deviations across different scales is misleading.
A poll of 400 likely voters finds that 52% support Candidate A, with a margin of error of ±4 percentage points at the 95% confidence level. A news anchor declares: 'This poll proves Candidate A will win.' Which of the following best identifies the flaw in this claim?
Answer: The confidence interval includes values below 50%, so a majority may not actually support Candidate A
The 95% confidence interval is 52% ± 4%, which spans from 48% to 56%. Because this interval includes values below 50%, it is statistically possible that fewer than half of all likely voters support Candidate A. The anchor's error is treating the point estimate (52%) as a certainty rather than recognizing the range of plausible true values the interval represents.
A streaming service surveys 500 users who watched at least one documentary in the past 30 days and finds that 78% of them prefer documentaries over other genres. The service concludes that 78% of all subscribers prefer documentaries. Which of the following best explains the primary flaw in this conclusion?
Answer: The sampling method systematically overrepresents users who already watch documentaries
This is a classic example of sampling bias. By restricting the survey to users who already watched a documentary, the service excluded the portion of subscribers who never watch documentaries — likely the group least inclined to prefer them. The sample is not representative of all subscribers, so the 78% figure cannot be generalized to the full population. Sample size (choice A) is irrelevant when the selection process is fundamentally biased.
A hospital compares two drugs across two independent trials. Drug A cured 80% of patients in Trial 1 and 75% in Trial 2. Drug B cured 70% of patients in Trial 1 and 65% in Trial 2. A researcher concludes Drug A is more effective overall. Which piece of additional information would most directly challenge this conclusion?
Answer: Whether Drug A treated far more patients in the lower-cure-rate trial while Drug B treated far more patients in the higher-cure-rate trial
This scenario describes Simpson's Paradox. Even though Drug A outperforms Drug B within each individual trial, the overall combined cure rate can reverse if the patient counts are unequal across trials. If Drug A treated mostly patients in Trial 2 (the harder trial, lower cure rates) while Drug B treated mostly patients in Trial 1 (the easier trial, higher cure rates), Drug B's overall weighted average could equal or exceed Drug A's. The within-group advantage does not guarantee an overall advantage when group sizes differ substantially.
A data analyst plots the number of fire trucks dispatched to fires (x-axis) against property damage in thousands of dollars (y-axis) and finds a strong positive correlation of r = 0.89. The analyst concludes: 'Dispatching more fire trucks causes more damage. The city should reduce fire truck deployments to save property.' What is the most significant error in this reasoning?
Answer: Both variables are likely driven by a third factor — fire severity — making the correlation spurious rather than causal
This is a textbook confounding variable (lurking variable) error. Larger, more severe fires independently cause both greater property damage and the dispatch of more fire trucks. Fire severity drives both variables simultaneously, creating a strong positive correlation that has nothing to do with trucks causing damage. Reducing truck deployments would not lower property damage — it would worsen it. Correlation, regardless of its strength, never establishes causation when a plausible confounding variable exists.