Free SAT Data Analysis & Statistics Questions and Answers 2 — Questions and Answers
Question 1: A survey samples 200 students from a school of 2,000 to estimate the average hours of sleep per night. The sample mean is 6.8 hours. What does this tell us about the school population?
- Every student in the school sleeps exactly 6.8 hours per night
- The school population is estimated to sleep an average of approximately 6.8 hours, with some uncertainty (Correct answer)
- The survey is invalid because it did not survey all 2,000 students
- No conclusions about the school can be drawn from a sample of only 200
Correct answer: The school population is estimated to sleep an average of approximately 6.8 hours, with some uncertainty
A representative sample allows inference about the population — the sample mean (6.8 hours) is an estimate of the population mean, with a margin of error depending on sample size and variability.
Statistical inference allows conclusions about a population based on a sample. With a properly drawn sample of 200 from 2,000 students: (1) The sample mean (6.8 hours) is an unbiased estimate of the population mean; (2) There is uncertainty — expressed as a margin of error or confidence interval; (3) The sample doesn't need to be the whole population to be informative. Common misconceptions: 'too small to be valid' (200 is actually a reasonable sample for this population) and 'everyone sleeps 6.8 hours' (the mean is an average, not an individual value). Sampling validity depends more on HOW the sample was selected (random vs. biased) than on its absolute size.
Question 2: A data set has the following values: 5, 7, 7, 8, 9, 10, 100. Which measure of center would be MOST misleading as a description of 'typical' values?
- Median
- Mean (Correct answer)
- Mode
- Range
Correct answer: Mean
The mean is pulled upward by the extreme outlier (100) to approximately 20.9, while most values cluster between 5 and 10. The median (8) better represents the 'typical' value. The mean is most misleading here.
Calculating for this data set {5, 7, 7, 8, 9, 10, 100}: Mean = (5+7+7+8+9+10+100)/7 = 146/7 ≈ 20.9. But 6 of 7 values are below 10 — so 20.9 is not typical of any value in the set. Median = 8 (middle value when sorted). Mode = 7 (most frequent). The mean is heavily influenced by the outlier (100) and is therefore the most misleading description of 'typical.' This is why median is preferred for skewed distributions (like income data with a few billionaires) — it is resistant to outliers.
Question 3: A graph shows that the number of library books checked out per month and the city's ice cream sales are both highest in summer. A researcher concludes that library visits cause ice cream purchases. This reasoning is flawed because:
- Summer data is always unreliable for research purposes
- Correlation does not imply causation — a third variable (summer season) likely drives both (Correct answer)
- The sample size of months is too small to draw conclusions
- Ice cream and library books are too unrelated to be compared
Correct answer: Correlation does not imply causation — a third variable (summer season) likely drives both
Both library visits and ice cream sales increase in summer — but summer (a third variable, or 'confound') is the underlying cause of both. This is classic spurious correlation driven by a confounding variable.
Spurious correlation occurs when two variables are correlated due to a common cause, not because either causes the other. In this case, warm summer weather increases both outdoor activities (leading to more library/community engagement) and ice cream consumption. The correlation between library visits and ice cream sales is 'spurious' — caused by the hidden third variable (season/weather). This type of confusion is extremely common in data analysis. The distinction between correlation and causation requires: controlled experiments (randomization), statistical control for confounding variables, or causal modeling with theoretical justification.
Question 4: A box plot shows a distribution with the median significantly closer to the lower quartile than the upper quartile. This indicates the distribution is:
- Approximately normal (symmetric)
- Right-skewed (positively skewed) — more values on the lower end with a tail extending higher (Correct answer)
- Left-skewed (negatively skewed) — a tail extending toward lower values
- Bimodal — having two peaks
Correct answer: Right-skewed (positively skewed) — more values on the lower end with a tail extending higher
When the median is closer to the lower quartile (Q1) and the upper 'box' extends further, the distribution has more scores in the lower range and a longer tail extending toward higher values — this is right (positive) skew.
In a box plot: The box spans Q1 to Q3 (the interquartile range). The median line divides the box. If the median is closer to Q1 than Q3: the upper portion of the box (Q2 to Q3) is longer — indicating more spread in the upper half. The distribution is right-skewed (positive skew) — most data clusters at lower values with a tail extending toward higher values. Recall: right skew means the tail points RIGHT (toward higher values); left skew means the tail points LEFT (toward lower values). Income distributions are typically right-skewed.
Question 5: A researcher conducts an experiment where participants are RANDOMLY ASSIGNED to treatment and control groups. This design feature primarily controls for:
- Measurement error in the data collection instruments
- Confounding variables that might otherwise create spurious results (Correct answer)
- Sample size limitations affecting statistical power
- Observer bias in data interpretation
Correct answer: Confounding variables that might otherwise create spurious results
Random assignment distributes both known and unknown confounding variables equally between groups — making the groups comparable so that any observed difference can be attributed to the treatment.
Random assignment (randomization) is the key feature that distinguishes a true experiment from an observational study. When participants are randomly assigned to conditions: (1) Both known confounders (age, gender, health status) AND unknown confounders are distributed equally across groups; (2) Any systematic difference observed between groups at the end can be attributed to the treatment — not to pre-existing differences; (3) This is the foundation of causal inference in experimental research. Without randomization, groups may differ in important ways that confound the results. Randomization is why clinical trials (RCTs) are considered the gold standard for establishing medical effectiveness.
Question 6: A student calculates that the correlation coefficient between study hours and test scores in their class is r = 0.85. What does this indicate?
- Studying 0.85 more hours will increase test scores by 1 point
- There is a strong positive relationship between study hours and test scores, though this doesn't prove causation (Correct answer)
- 85% of students who study will pass the test
- Test scores increase by 85% for each additional study hour
Correct answer: There is a strong positive relationship between study hours and test scores, though this doesn't prove causation
A correlation coefficient of 0.85 indicates a strong positive linear relationship (as study hours increase, scores tend to increase), but correlation alone does not prove that studying CAUSES higher scores.
Correlation coefficients (r) range from -1 to +1: r = 1: perfect positive correlation; r = 0.85: strong positive correlation; r = 0: no linear relationship; r = -0.85: strong negative correlation; r = -1: perfect negative correlation. A correlation of 0.85 means: when study hours are above average, test scores tend to be above average (strong positive tendency). However: (1) This doesn't prove causation — students who study more may also have other advantageous characteristics; (2) The coefficient doesn't mean 'percentage improvement'; (3) It describes a tendency across many students, not a predictable individual change.
A survey samples 200 students from a school of 2,000 to estimate the average hours of sleep per night.
The sample mean is 6.8 hours.
What does this tell us about the school population?