CBT Psychometric Principles and Analysis 3 — Questions and Answers
Question 1: A passing score on a licensure exam is set by having subject matter experts rate the minimum competency required to answer each item correctly. This method is called:
- Norm-referenced standard setting
- The Angoff method (Correct answer)
- Equipercentile equating
- Rasch scaling
Correct answer: The Angoff method
The Angoff method asks judges to estimate the probability that a minimally competent examinee would answer each item correctly, and averages these ratings to set the cut score.
Question 2: In generalizability theory, the 'universe score' is conceptually analogous to which classical test theory concept?
- Standard error of measurement
- True score (Correct answer)
- Observed score
- Reliability coefficient
Correct answer: True score
The universe score in G-theory represents the expected score over all conditions of measurement, paralleling the true score in classical test theory.
Question 3: Which type of score interpretation compares an examinee's performance to a defined content domain rather than to other examinees?
- Norm-referenced interpretation
- Criterion-referenced interpretation (Correct answer)
- Ipsative interpretation
- Standardized interpretation
Correct answer: Criterion-referenced interpretation
Criterion-referenced interpretation evaluates mastery of a defined domain, reporting what an examinee can or cannot do rather than comparing them to a normative group.
Question 4: An examinee's score on a cognitive test is 115. The test has a reliability of 0.84 and a mean of 100. What concept explains why the examinee's true score estimate should be pulled toward the mean?
- Ceiling effect
- Regression to the mean (Correct answer)
- Halo effect
- Response bias
Correct answer: Regression to the mean
Regression to the mean occurs because observed scores contain measurement error; true score estimates are weighted averages of the observed score and the group mean, pulled toward the mean by unreliability.
Question 5: What does a negative point-biserial correlation for a test item most likely indicate?
- The item is highly discriminating
- Higher-scoring examinees tend to answer the item incorrectly (Correct answer)
- The item has a low difficulty level
- The item has too many response options
Correct answer: Higher-scoring examinees tend to answer the item incorrectly
A negative point-biserial means low total scorers pass the item more than high scorers, suggesting the item is flawed, miskeyed, or measures something different from the rest of the test.
Question 6: In the context of differential item functioning (DIF), which condition is labeled 'uniform DIF'?
- One group consistently outperforms another across all ability levels (Correct answer)
- The advantage for one group reverses at different ability levels
- Both groups have equal mean scores but different variances
- DIF is present only on speeded items
Correct answer: One group consistently outperforms another across all ability levels
Uniform DIF occurs when one group has a consistently higher probability of correct response across all ability levels, with no interaction between group membership and ability.
Question 7: Which of the following indices is specifically designed to evaluate model-data fit for individual items under Item Response Theory?
- Cronbach's alpha
- Root mean square error of approximation (RMSEA)
- Outfit mean square statistic (Correct answer)
- Coefficient omega
Correct answer: Outfit mean square statistic
The outfit mean square statistic (and its counterpart infit) measures discrepancy between observed and model-expected responses for individual items, flagging item misfit in IRT.
A passing score on a licensure exam is set by having subject matter experts rate the minimum competency required to answer each item correctly.
This method is called: