← All Bluebook SAT Test Flashcard Decks

Data Analysis & Statistics Flashcards

6 cards from real Bluebook SAT Test practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Data Analysis & Statistics flashcards as text
  1. A researcher collects data on study hours and exam scores for 80 students. The correlation coefficient between study hours and exam scores is r = 0.87. The researcher concludes that for every additional hour studied, a student's exam score increases by 0.87 points. Which of the following best describes the error in the researcher's reasoning?

    Answer: The correlation coefficient only describes the strength and direction of a linear relationship, not the rate of change in scores per hour.

    The correlation coefficient r = 0.87 indicates a strong positive linear relationship between study hours and exam scores, but it does NOT describe the slope (rate of change). The slope of the regression line — not r — tells you how many points a score changes per additional hour studied. Confusing r with the slope is a classic misinterpretation.

  2. A data set of 9 values has a mean of 20 and a median of 17. When a 10th value is added, the mean increases to 21. What is the new median if the 10th value added is 30?

    Answer: 18.5

    The original 9 values sum to 9 × 20 = 180. The new mean is 21 with 10 values, so the new sum is 210, confirming the added value is 210 − 180 = 30. Now there are 10 values. The original median of 17 was the 5th value. With 30 added (which is larger than the current median), the sorted order shifts so the new median is the average of the 5th and 6th values. Since the median was 17 (the 5th of 9), the 5th and 6th values of the new sorted 10-value set are 17 and 20 (the next value above median), giving a new median of (17 + 20)/2 = 18.5.

  3. In a study, Group A received a new tutoring program and Group B did not. After 6 weeks, Group A's average test score improved by 15 points while Group B's average improved by 11 points. The researchers randomly assigned students to groups. Which of the following is the most justified conclusion?

    Answer: The tutoring program likely caused the additional 4-point improvement in Group A.

    Because students were randomly assigned to groups, this is a randomized controlled experiment, which does allow causal conclusions. The most justified conclusion is that the tutoring program likely caused the additional improvement. Choice B introduces an unsupported confounding variable. Choice C overgeneralizes beyond the study. Choice D misuses the word 'proves' — statistical significance requires a formal test, and the 4-point difference alone doesn't establish it.

  4. A scatterplot shows data for 12 cities plotting annual rainfall (x-axis, in inches) versus annual car wash revenue (y-axis, in thousands of dollars). The line of best fit has a negative slope. A city with 45 inches of annual rainfall is an outlier that lies far above the line of best fit. Which of the following best describes this outlier city?

    Answer: It has much higher car wash revenue than predicted by its rainfall level.

    An outlier that lies 'far above the line of best fit' means its actual y-value (car wash revenue) is much greater than what the regression line predicts for its x-value (45 inches of rainfall). Since the line has a negative slope, cities with 45 inches of rain would be predicted to have relatively low revenue — but this city's revenue is much higher than predicted, placing it above the line.

  5. A survey of 400 randomly selected adults in a city asked whether they support a new transit policy. The results showed 55% in support. The margin of error was reported as ±4.9% at a 95% confidence level. A city official claims: 'We are 95% confident that between 50.1% and 59.9% of all adults in the city support this policy.' Which of the following identifies a flaw in this interpretation?

    Answer: The confidence interval describes where the true population proportion likely falls, not a probability that it is in that range after the interval is computed.

    The official's phrasing that 'we are 95% confident the true proportion is between 50.1% and 59.9%' is a common but subtle misinterpretation. Once the interval is computed, the true population proportion either is or is not in that range — it's not a 95% probability. Correctly stated, 95% confidence means that if this process were repeated many times, 95% of the resulting intervals would capture the true proportion. The other choices misrepresent how margins of error and sample sizes function.

  6. The table below shows the distribution of scores on a 10-point quiz for two classes: Class A: scores {4, 5, 6, 7, 7, 8, 8, 8, 9, 10} Class B: scores {1, 2, 4, 7, 7, 8, 9, 10, 10, 10} Which of the following correctly compares the two distributions?

    Answer: Class A and Class B have the same median, but Class B has a larger standard deviation.

    Both Class A and Class B have 10 scores. Class A sorted: {4,5,6,7,7,8,8,8,9,10} — median = (7+8)/2 = 7.5. Class B sorted: {1,2,4,7,7,8,9,10,10,10} — median = (7+8)/2 = 7.5. The medians are equal. Class A's scores cluster tightly around 7–8, while Class B's scores are more spread out (ranging from 1 to 10 with values at extremes), giving Class B a larger standard deviation. The mean of Class A = 72/10 = 7.2; mean of Class B = 68/10 = 6.8, so Class A has the higher mean — eliminating choices C and D.