ETS Validity and Reliability 3 — Questions and Answers
Question 1: Cronbach's alpha is most appropriate for tests where items are scored:
- Right or wrong only
- On a polytomous (multi-point) scale (Correct answer)
- By multiple raters
- At two separate time points
Correct answer: On a polytomous (multi-point) scale
Cronbach's alpha generalizes KR-20 to items with more than two score categories, such as Likert-scale responses.
Question 2: A validity coefficient of r = 0.40 between an entrance exam and GPA is considered:
- Too low to be useful in practice
- Moderate and practically useful for selection decisions (Correct answer)
- Perfect and uncommon in social science research
- A sign of poor test construction
Correct answer: Moderate and practically useful for selection decisions
In educational and employment testing, validity coefficients around 0.30–0.50 are considered moderate and practically useful.
Question 3: Which of the following would BEST demonstrate discriminant validity?
- A new anxiety scale correlates highly with a depression scale
- A new anxiety scale correlates low with an unrelated measure like shoe size (Correct answer)
- A new anxiety scale has high internal consistency
- A new anxiety scale predicts therapy outcomes well
Correct answer: A new anxiety scale correlates low with an unrelated measure like shoe size
Discriminant validity is supported when a measure does NOT correlate strongly with measures of unrelated constructs.
Question 4: The Spearman-Brown formula is used to:
- Correct a split-half reliability estimate for test length (Correct answer)
- Convert raw scores to percentile ranks
- Estimate the standard error of measurement
- Calculate concurrent validity coefficients
Correct answer: Correct a split-half reliability estimate for test length
The Spearman-Brown prophecy formula corrects a split-half correlation (based on half the test) to estimate reliability for the full test length.
Question 5: A test developer wants to ensure a reading test covers all major subskills of reading comprehension. This concern relates to:
- Consequential validity
- Content validity (Correct answer)
- Predictive validity
- Convergent validity
Correct answer: Content validity
Content validity reflects whether test items adequately and representatively sample the domain of skills or knowledge the test intends to measure.
Question 6: Generalizability theory extends classical test theory by allowing researchers to examine:
- Only random error sources
- Multiple sources of measurement error simultaneously (Correct answer)
- Only inter-rater disagreement
- Face validity across populations
Correct answer: Multiple sources of measurement error simultaneously
Generalizability theory (G-theory) partitions score variance into multiple facets (raters, items, occasions), providing a richer picture of measurement error.
Question 7: If a test's standard error of measurement (SEM) is large, which statement is TRUE?
- Individual scores can be interpreted with high confidence
- The test has high content validity
- Score precision is low and confidence intervals are wide (Correct answer)
- Criterion validity is necessarily high
Correct answer: Score precision is low and confidence intervals are wide
A large SEM means individual observed scores vary greatly around the true score, reducing precision and widening score confidence intervals.
Cronbach's alpha is most appropriate for tests where items are scored: