ATP Research & Evidence-Based Practice 2 — Questions and Answers
Question 1: Which statistical concept describes the degree to which a test score reflects true ability rather than measurement error?
- Standard error of measurement
- Reliability coefficient (Correct answer)
- Validity coefficient
- Item discrimination index
Correct answer: Reliability coefficient
The reliability coefficient quantifies the proportion of observed score variance attributable to true score variance, directly reflecting measurement precision.
Question 2: In Item Response Theory, the 'b' parameter of a dichotomous item represents:
- Item discrimination
- Item difficulty (location) (Correct answer)
- Pseudo-guessing probability
- Item information
Correct answer: Item difficulty (location)
The b parameter in IRT is the difficulty or location parameter, indicating the ability level at which a test-taker has a 50% probability of answering correctly (in the 1PL/2PL models).
Question 3: A researcher finds that exam scores predict job performance equally well for male and female applicants. This demonstrates:
- Content validity
- Predictive bias absence (Correct answer)
- Construct equivalence
- Criterion contamination
Correct answer: Predictive bias absence
When regression lines for predicting a criterion from test scores do not differ significantly across subgroups, the test shows no predictive bias.
Question 4: Which research design provides the strongest evidence for causal inference in educational testing?
- Cross-sectional survey
- Randomized controlled trial (Correct answer)
- Correlational study
- Case study
Correct answer: Randomized controlled trial
Randomized controlled trials allow causal conclusions because random assignment controls for confounding variables.
Question 5: The Standards for Educational and Psychological Testing defines 'consequential validity' as primarily concerned with:
- Whether items are worded clearly
- Intended and unintended social consequences of test use (Correct answer)
- The stability of scores over time
- The match between test blueprint and content
Correct answer: Intended and unintended social consequences of test use
Consequential validity examines the value implications and social consequences—both intended and unintended—that result from test score interpretations and uses.
Question 6: Differential item functioning (DIF) analysis is primarily used to detect:
- Items that are too easy for the reference group
- Items that perform differently across groups after matching on ability (Correct answer)
- Items with low point-biserial correlations
- Items that exceed target difficulty levels
Correct answer: Items that perform differently across groups after matching on ability
DIF identifies items where examinees of equal ability from different groups have systematically different probabilities of a correct response, suggesting potential bias.
Question 7: Which approach to standard setting requires panelists to estimate the probability that a minimally competent candidate would answer each item correctly?
- Angoff method (Correct answer)
- Bookmark method
- Contrasting groups method
- Body of work method
Correct answer: Angoff method
The Angoff method asks panelists to estimate the proportion of minimally competent examinees expected to answer each item correctly, then averages these estimates.
Which statistical concept describes the degree to which a test score reflects true ability rather than measurement error?