ATP Research & Evidence-Based Practice 3 — Questions and Answers
Question 1: Which type of validity evidence is gathered by examining the internal structure of a test through factor analysis?
- Evidence based on relations to other variables
- Evidence based on internal structure (Correct answer)
- Evidence based on test content
- Evidence based on response processes
Correct answer: Evidence based on internal structure
The Standards recognize internal structure as a source of validity evidence; factor analysis examines whether item relationships align with the intended construct.
Question 2: In a meta-analysis of test validity studies, the process of correcting for attenuation accounts for:
- Sampling error only
- Unreliability in predictor and criterion measures (Correct answer)
- Range restriction only
- Publication bias
Correct answer: Unreliability in predictor and criterion measures
Correction for attenuation adjusts validity coefficients upward to estimate what the true relationship would be if both predictor and criterion were measured without error.
Question 3: A test developer uses think-aloud protocols during pilot testing. This method primarily provides evidence of:
- Criterion-related validity
- Response process validity (Correct answer)
- Content validity
- Predictive validity
Correct answer: Response process validity
Think-aloud protocols reveal the cognitive processes examinees use when answering items, providing evidence that response processes align with the intended construct.
Question 4: The Spearman-Brown prophecy formula is used to estimate:
- How reliability changes when test length is altered (Correct answer)
- The standard error of the mean
- Inter-rater reliability for open-ended items
- Effect size in group comparisons
Correct answer: How reliability changes when test length is altered
The Spearman-Brown formula predicts how reliability will increase or decrease as test length is multiplied by a given factor.
Question 5: In computerized adaptive testing (CAT), item selection is primarily driven by:
- Randomization to prevent item exposure
- Maximizing information at the current ability estimate (Correct answer)
- Matching item difficulty to the passing standard
- Ensuring content blueprint coverage first
Correct answer: Maximizing information at the current ability estimate
CAT algorithms select items that provide maximum statistical information (minimal measurement error) at the examinee's current estimated ability level.
Question 6: Which type of score reporting gives examinees feedback tied directly to specific skill domains rather than a single total score?
- Norm-referenced reporting
- Subscore reporting (Correct answer)
- Standard score reporting
- Percentile rank reporting
Correct answer: Subscore reporting
Subscore reporting disaggregates overall performance into domain-specific scores, providing actionable diagnostic information about specific skill areas.
Question 7: Criterion contamination occurs when:
- The criterion measure is unreliable
- Knowledge of test scores influences criterion ratings (Correct answer)
- The criterion is collected before the predictor
- The criterion measure lacks face validity
Correct answer: Knowledge of test scores influences criterion ratings
Criterion contamination happens when raters or supervisors providing criterion data are aware of examinees' test scores, inflating the apparent validity coefficient.
Which type of validity evidence is gathered by examining the internal structure of a test through factor analysis?