ATP Research & Evidence-Based Practice 5 — Questions and Answers
Question 1: Which approach to validity argument, articulated by Messick, emphasizes that validity is a unitary concept focused on:
- Multiple independent validity coefficients
- The appropriateness of inferences and actions based on test scores (Correct answer)
- Reliability of scores across occasions
- The number of items per content domain
Correct answer: The appropriateness of inferences and actions based on test scores
Messick's unified conception holds that validity is a single integrated judgment about the degree to which evidence supports the intended interpretations and uses of test scores.
Question 2: A test developer discovers that a high-stakes certification exam has a significant adverse impact ratio below 4/5 for one ethnic group. The NEXT recommended step is to:
- Remove the affected items immediately
- Investigate whether the score differences reflect construct-relevant or irrelevant factors (Correct answer)
- Lower the passing standard to equalize pass rates
- Suspend test administration pending a full retraction
Correct answer: Investigate whether the score differences reflect construct-relevant or irrelevant factors
The 4/5 rule triggers further investigation, not automatic action; the critical question is whether group differences stem from real construct-related factors or test bias.
Question 3: Which reliability approach is most appropriate for assessing inter-rater agreement on a constructed-response item with ordinal scoring rubrics?
- Pearson product-moment correlation
- Weighted kappa (Correct answer)
- Coefficient alpha
- Test-retest correlation
Correct answer: Weighted kappa
Weighted kappa accounts for ordinal scoring by penalizing disagreements proportionally to their distance, making it more appropriate than simple kappa or Pearson r for rubric-scored responses.
Question 4: Generalizability theory extends classical test theory by partitioning score variance into multiple facets, such as:
- True score, error, and guessing components
- Person, item, rater, and occasion effects (Correct answer)
- Difficulty, discrimination, and pseudo-chance parameters
- Content, cognitive, and affective domains
Correct answer: Person, item, rater, and occasion effects
G-theory designs allow simultaneous estimation of variance components attributable to persons, items, raters, occasions, and their interactions.
Question 5: A licensing board wants to set a defensible passing score. Which of the following best strengthens the legal defensibility of the standard-setting process?
- Using the 50th percentile as the passing score
- Documenting the qualifications of panelists and the procedural steps followed (Correct answer)
- Selecting the cut score that maximizes pass rates
- Relying solely on normative data from prior test administrations
Correct answer: Documenting the qualifications of panelists and the procedural steps followed
Legal defensibility depends on transparent documentation of panelist qualifications, training, procedures, and rationale, demonstrating due diligence in the standard-setting process.
Question 6: Which ethical obligation do test publishers have when research reveals that an operational test measure may be functioning as intended for some populations but not others?
- Maintain the status quo to avoid disrupting operational programs
- Disclose findings, investigate causes, and take corrective action (Correct answer)
- Restrict publication of the research to avoid public concern
- Immediately retire the test and refund all fees
Correct answer: Disclose findings, investigate causes, and take corrective action
Professional and ethical standards require publishers to disclose limitations, investigate differential functioning, and implement corrective measures to protect test-taker welfare.
Question 7: Which research synthesis method allows quantitative aggregation of effect sizes across multiple independent validity studies?
- Narrative literature review
- Meta-analysis (Correct answer)
- Systematic review without statistics
- Expert consensus panel
Correct answer: Meta-analysis
Meta-analysis statistically combines effect sizes (e.g., validity coefficients) from multiple studies, providing a more precise and generalizable estimate than any single study.
Which approach to validity argument, articulated by Messick, emphasizes that validity is a unitary concept focused on: