ATP Research & Evidence-Based Practice 4 — Questions and Answers
Question 1: Which reliability estimation method is most appropriate when test items measure multiple distinct dimensions?
- Coefficient alpha
- Alternate forms reliability (Correct answer)
- Split-half with Spearman-Brown correction
- Test-retest reliability
Correct answer: Alternate forms reliability
Coefficient alpha assumes unidimensionality and underestimates reliability for multidimensional tests; alternate forms reliability does not carry this assumption.
Question 2: In evidence-based assessment, an incremental validity study demonstrates that a new test:
- Has higher reliability than existing measures
- Explains additional variance in a criterion beyond existing predictors (Correct answer)
- Is more cost-effective than current tests
- Produces smaller subgroup score differences
Correct answer: Explains additional variance in a criterion beyond existing predictors
Incremental validity shows that a test adds predictive power beyond what can be explained by already-available information, justifying its practical use.
Question 3: The 'table of specifications' in test development primarily serves to ensure:
- Items are written at appropriate reading levels
- Content coverage aligns with the intended domain blueprint (Correct answer)
- All items have similar difficulty levels
- The test meets time-limit requirements
Correct answer: Content coverage aligns with the intended domain blueprint
A table of specifications maps items to content areas and cognitive levels, ensuring the test adequately and proportionally samples the target knowledge domain.
Question 4: Which statistical index is used to evaluate how well a set of items in a CAT fits the IRT model for an individual examinee?
- Infit and outfit mean square statistics (Correct answer)
- Point-biserial correlation
- Cronbach's alpha
- Chi-square goodness-of-fit across all items
Correct answer: Infit and outfit mean square statistics
Infit and outfit statistics from Rasch/IRT analysis quantify the degree to which individual item responses conform to model expectations for that examinee.
Question 5: A test publisher conducts a predictive validity study with a two-year follow-up. Which threat to internal validity is most relevant?
- Testing effects
- Attrition/mortality (Correct answer)
- Instrumentation changes
- Maturation
Correct answer: Attrition/mortality
Over a two-year period, differential attrition—where certain types of participants drop out—can systematically bias the sample and distort the observed validity coefficient.
Question 6: When applying the Bookmark standard-setting method, ordered item booklets are arranged by:
- Alphabetical order of content standards
- Increasing probability of a correct response for a typical examinee (Correct answer)
- Decreasing item discrimination values
- Increasing item length and complexity
Correct answer: Increasing probability of a correct response for a typical examinee
In the Bookmark method, items are sorted in order of increasing difficulty (based on IRT-estimated probability of correct response), and panelists place a bookmark where minimally competent candidates would likely stop succeeding.
Question 7: Which principle of evidence-based practice requires that research findings be applied in ways that consider local context and examinee characteristics?
- Statistical significance alone guides decisions
- Generalizability of findings to the local population (Correct answer)
- Replication of the exact study design
- Use of the largest available normative sample
Correct answer: Generalizability of findings to the local population
Evidence-based practice requires practitioners to judge whether research evidence generalizes appropriately to their specific context, population, and purpose.
Which reliability estimation method is most appropriate when test items measure multiple distinct dimensions?