Data Interpretation & Analysis Flashcards
6 cards from real NBT practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Data Interpretation & Analysis flashcards as text
A researcher plots the residuals from a linear regression of study hours vs. test scores. The residual plot shows a clear U-shaped (parabolic) pattern. What does this most likely indicate?
Answer: The relationship between the variables is non-linear, and a linear model is inappropriate
A U-shaped (curved) pattern in a residual plot is a classic diagnostic sign that the relationship between the predictor and response variable is non-linear. Linear regression assumes a linear relationship; a systematic curve in the residuals means the model is systematically under- or over-predicting, indicating a quadratic or other non-linear model would be more appropriate. Heteroscedasticity would appear as a fan-shaped spread, not a curve.
Two datasets each have a mean of 50. Dataset A has a standard deviation of 2 and Dataset B has a standard deviation of 12. A value of 56 is drawn from one of the datasets. From which dataset is this value more surprising, and why?
Answer: Dataset A, because 56 is 3 standard deviations above the mean in Dataset A but only 0.5 standard deviations above in Dataset B
Surprise (or unusualness) is measured by how many standard deviations a value lies from the mean — the z-score. For Dataset A: z = (56−50)/2 = 3, placing 56 far into the tail. For Dataset B: z = (56−50)/12 ≈ 0.5, which is very close to the mean. A z-score of 3 is far more extreme than 0.5, making 56 much more surprising in Dataset A despite both having the same mean.
The table below summarizes annual income (in thousands of rands) for two towns: Town X: Median = R320k, Mean = R510k, Mode = R280k Town Y: Median = R315k, Mean = R318k, Mode = R310k A policy analyst wants to describe the 'typical' resident's income in each town. Which measures and conclusion are most analytically sound?
Answer: Use the median for both towns; Town X's distribution is heavily right-skewed by high earners, while Town Y's incomes are roughly symmetric
In Town X, the mean (R510k) is dramatically higher than the median (R320k) and mode (R280k), which is a strong indicator of right skew — a small number of very high earners are pulling the mean up. The median is resistant to outliers and better represents the 'typical' resident. In Town Y, the mean, median, and mode are nearly identical (~R315–318k), indicating a roughly symmetric distribution where the mean is also appropriate — but the median remains the safest choice for 'typical' income comparisons across both towns for consistency.
A study reports a correlation coefficient of r = 0.82 between ice cream sales and drowning incidents across 24 months. The p-value is 0.001. A journalist concludes that eating ice cream increases drowning risk. Which critique is most statistically rigorous?
Answer: The correlation is likely explained by a confounding variable (hot weather), and correlation never establishes causation regardless of statistical significance
This is a classic example of a spurious correlation driven by a confounding variable — hot weather simultaneously increases both ice cream consumption and swimming (which raises drowning risk). Statistical significance (low p-value) only tells us the correlation is unlikely due to chance; it says nothing about causation. A low p-value with r = 0.82 simply means the correlation is real and consistent — not that one variable causes the other. Establishing causation requires controlled experimental design, not just observational correlation.
A box-and-whisker plot for a dataset shows: minimum = 10, Q1 = 25, median = 40, Q3 = 60, maximum = 130. Using the standard IQR fence method, which values should be flagged as potential outliers?
Answer: Only 130, because the upper fence is Q3 + 1.5×IQR = 60 + 52.5 = 112.5, and 130 exceeds it
The IQR = Q3 − Q1 = 60 − 25 = 35. The upper fence = Q3 + 1.5×IQR = 60 + 52.5 = 112.5. The lower fence = Q1 − 1.5×IQR = 25 − 52.5 = −27.5. The minimum value of 10 is above the lower fence (−27.5), so it is not an outlier. The maximum value of 130 exceeds the upper fence of 112.5, so it IS flagged as a potential outlier. The criterion is based on IQR fences, not on visual extremity or ratio to the median.
A survey of 400 university students finds that 60% prefer online learning. The 95% confidence interval is reported as (55.2%, 64.8%). A dean interprets this as: 'There is a 95% probability that the true population proportion falls between 55.2% and 64.8%.' What is wrong with this interpretation?
Answer: The interval is either correct or not — the 95% refers to the long-run reliability of the method, not the probability that this specific interval contains the true proportion
This is one of the most common misinterpretations of confidence intervals. Once a specific interval is computed, the true population proportion either falls within it or it doesn't — there is no probabilistic 'chance' about that particular interval. The correct interpretation is: if we repeated this sampling process many times and constructed a 95% CI each time, approximately 95% of those intervals would contain the true proportion. The 95% describes the reliability of the *procedure*, not the probability that *this specific computed interval* is correct.