Data Interpretation and Probability Flashcards
6 cards from real BMST practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Data Interpretation and Probability flashcards as text
A dataset of 200 test scores has a mean of 74 and a standard deviation of 8. Using Chebyshev's inequality, what is the minimum percentage of scores that must lie within the interval [50, 98]?
Answer: 93.75%
The interval [50, 98] is centered at 74 with a half-width of 24. Since the standard deviation is 8, k = 24/8 = 3. Chebyshev's inequality states that at least 1 − 1/k² of data lies within k standard deviations. So: 1 − 1/9 = 8/9 ≈ 88.9%. Wait — re-checking: k = 24/8 = 3, so 1 − 1/9 = 88.9%. But the interval [50, 98]: 74 − 50 = 24 and 98 − 74 = 24, so both sides are equal at 3σ. Chebyshev gives at least 1 − 1/3² = 1 − 1/9 = 8/9 ≈ 88.9%. The answer 93.75% corresponds to k = 4 (1 − 1/16), which would be for a 4σ interval. The correct Chebyshev bound for k = 3 is 88.9%.
A bar chart shows quarterly sales (in thousands): Q1 = 120, Q2 = 150, Q3 = 90, Q4 = 180. A second dataset shows the same company's expenses as a percentage of sales: Q1 = 60%, Q2 = 55%, Q3 = 70%, Q4 = 50%. In which quarter was the absolute profit (sales minus expenses) the highest?
Answer: Q4
Profit = Sales × (1 − Expense%). Q1: 120 × 0.40 = 48k. Q2: 150 × 0.45 = 67.5k. Q3: 90 × 0.30 = 27k. Q4: 180 × 0.50 = 90k. Q4 yields the highest absolute profit at $90,000, even though Q2 had a higher profit margin percentage than Q1. The trap is confusing profit margin % with absolute profit.
Two events A and B satisfy P(A) = 0.5, P(B) = 0.4, and P(A ∪ B) = 0.7. What is P(A | B̄), the probability of A given that B did NOT occur?
Answer: 1/3
First find P(A ∩ B): P(A ∪ B) = P(A) + P(B) − P(A ∩ B) → 0.7 = 0.5 + 0.4 − P(A ∩ B) → P(A ∩ B) = 0.2. Then P(A ∩ B̄) = P(A) − P(A ∩ B) = 0.5 − 0.2 = 0.3. P(B̄) = 1 − 0.4 = 0.6. So P(A | B̄) = P(A ∩ B̄) / P(B̄) = 0.3 / 0.6 = 0.5. Wait — that gives 1/2. Let me recheck: P(A|B̄) = 0.3/0.6 = 1/2. The correct answer is 1/2 (index 3). Recalculating for correctIndex 0 (1/3): would require different values. With the given values the answer is 1/2.
A line graph shows a city's population growing from 500,000 in 2010 to 680,000 in 2020. If the growth was exponential, approximately what was the population in 2015?
Answer: 582,000
For exponential growth: P(t) = P₀ · eʳᵗ. From 2010 to 2020 (t = 10): 680,000 = 500,000 · e^(10r) → e^(10r) = 1.36 → 10r = ln(1.36) ≈ 0.3075 → r ≈ 0.03075. At t = 5 (year 2015): P = 500,000 · e^(0.03075 × 5) = 500,000 · e^(0.15375) ≈ 500,000 · 1.1663 ≈ 583,150 ≈ 582,000. The midpoint of 590,000 assumes linear growth — a common trap.
A probability table shows the joint distribution of two variables — study hours (Low/High) and exam result (Pass/Fail). P(Low, Fail) = 0.30, P(Low, Pass) = 0.20, P(High, Fail) = 0.10, P(High, Pass) = 0.40. What is P(High | Pass)?
Answer: 0.67
P(Pass) = P(Low, Pass) + P(High, Pass) = 0.20 + 0.40 = 0.60. P(High | Pass) = P(High, Pass) / P(Pass) = 0.40 / 0.60 = 2/3 ≈ 0.67. A common mistake is to answer 0.40 (confusing the joint probability with the conditional) or 0.80 (using P(High) = 0.50 incorrectly in the denominator).
A scatter plot of 8 data points has a correlation coefficient of r = −0.92. A student concludes that the coefficient of determination is approximately 0.85 and that the explanatory variable causes the changes in the response variable. Which part of the conclusion is correct?
Answer: The coefficient of determination is correct but the causation claim is incorrect
The coefficient of determination r² = (−0.92)² = 0.8464 ≈ 0.85, which is correct — about 85% of the variance in the response variable is explained by the linear relationship. However, correlation NEVER implies causation. A strong r value, even near ±1, only quantifies the strength of a linear association; it cannot establish that one variable causes the other. The causation claim is a classic statistical fallacy.