ETS Standardized Test Design 3 — Questions and Answers
Question 1: Which approach to standard-setting asks panelists to estimate the probability that a 'minimally competent' candidate would answer each item correctly?
- Angoff method (Correct answer)
- Bookmark method
- Body of Work method
- Contrasting Groups method
Correct answer: Angoff method
The Angoff method asks panelists to estimate the proportion of minimally competent candidates who would answer each item correctly, and those estimates are averaged to set the cut score.
Question 2: When a test measures a single underlying construct, it is said to demonstrate:
- Convergent validity
- Unidimensionality (Correct answer)
- Concurrent validity
- Factorial invariance
Correct answer: Unidimensionality
Unidimensionality means that performance on test items is primarily driven by one latent trait, which is a key assumption of many IRT models.
Question 3: A 'speeded test' is one in which:
- Items increase in difficulty as the test progresses
- Time limits prevent many examinees from attempting all items (Correct answer)
- Questions are drawn randomly from an item bank
- Scores are reported as percentile ranks
Correct answer: Time limits prevent many examinees from attempting all items
In a speeded test, the time limit is a major factor—examinees who work more quickly will attempt and complete more items, affecting score interpretation.
Question 4: Which measure of reliability estimates consistency across two different forms of the same test given simultaneously?
- Test-retest reliability
- Internal consistency
- Alternate-forms reliability (Correct answer)
- Inter-rater reliability
Correct answer: Alternate-forms reliability
Alternate-forms reliability (parallel-forms reliability) assesses consistency between two equivalent test forms administered in close succession.
Question 5: In a norm-referenced test, a student's score is interpreted by:
- Comparing performance to a predefined mastery standard
- Evaluating how many items were answered correctly
- Ranking performance relative to a defined reference group (Correct answer)
- Assessing competency in each content strand separately
Correct answer: Ranking performance relative to a defined reference group
Norm-referenced interpretation compares an examinee's score to the performance of a normative sample, typically expressed as percentile ranks or standard scores.
Question 6: Which type of validity evidence examines whether scores on a new test correlate with scores on an established measure of the same construct taken at the same time?
- Predictive validity
- Concurrent validity (Correct answer)
- Face validity
- Incremental validity
Correct answer: Concurrent validity
Concurrent validity is established when the new test's scores correlate with those of an existing criterion measure administered simultaneously.
Question 7: The standard error of measurement (SEM) is most directly used to:
- Estimate score reliability coefficients
- Construct confidence intervals around individual scores (Correct answer)
- Evaluate item discrimination
- Calculate percentile ranks
Correct answer: Construct confidence intervals around individual scores
The SEM quantifies measurement error and is used to build confidence intervals around an observed score to estimate where the true score likely falls.
Which approach to standard-setting asks panelists to estimate the probability that a 'minimally competent' candidate would answer each item correctly?