CBT Scoring Models and Reporting 5 — Questions and Answers
Question 1: What is the primary advantage of using IRT over Classical Test Theory (CTT) for CBT item calibration?
- IRT requires fewer examinees to calibrate items
- IRT item statistics are independent of the sample used for calibration (Correct answer)
- IRT eliminates the need for standard setting
- IRT scores are always higher than CTT scores
Correct answer: IRT item statistics are independent of the sample used for calibration
IRT item parameters are theoretically sample-independent (invariant), meaning item difficulty and discrimination estimates remain stable across different examinee groups.
Question 2: In which situation would a testing program most appropriately use norm-referenced scoring rather than criterion-referenced scoring?
- Licensing exams that certify minimum competency
- Selection exams used to rank applicants for limited spots (Correct answer)
- Formative classroom assessments checking mastery
- Clinical certification exams with a fixed passing standard
Correct answer: Selection exams used to rank applicants for limited spots
Norm-referenced scoring is appropriate when the goal is ranking and selecting a fixed number of candidates, such as admissions or competitive job selection.
Question 3: What does 'conditional standard error of measurement' (CSEM) provide that a single overall SEM cannot?
- A single precision estimate applicable to all scores
- Precision estimates that vary at different points along the score scale (Correct answer)
- A method for setting the cut score
- A measure of test validity
Correct answer: Precision estimates that vary at different points along the score scale
CSEM shows how measurement precision varies at each ability level, which is particularly valuable for CAT where precision near the cut score is most critical.
Question 4: A testing program reports 'classification consistency.' This statistic measures:
- Whether the test content matches the job analysis
- The likelihood that examinees receive the same pass/fail decision on retesting (Correct answer)
- How closely scores match criterion measures
- The correlation between test scores and academic grades
Correct answer: The likelihood that examinees receive the same pass/fail decision on retesting
Classification consistency estimates the probability that an examinee would receive the same pass or fail decision if they retested under identical conditions.
Question 5: Which phenomenon explains why a group's average score on a CBT tends to rise over multiple years without any actual improvement in knowledge?
- Regression to the mean
- Teach-to-the-test and score inflation (Correct answer)
- Differential item functioning
- Vertical scaling drift
Correct answer: Teach-to-the-test and score inflation
When item content becomes familiar through coaching and test-prep, scores can rise due to exposure rather than genuine competency gains, a phenomenon known as score inflation.
Question 6: What is the role of the 'anchor items' (common items) in the Non-Equivalent Anchor Test (NEAT) equating design?
- They are excluded from scoring to prevent bias
- They link the two test forms by providing a common statistical thread across different examinee groups (Correct answer)
- They are the hardest items on the exam
- They serve as practice items not counted in the final score
Correct answer: They link the two test forms by providing a common statistical thread across different examinee groups
Anchor items appear on both test forms and allow score equating by statistically linking the forms even when different groups of examinees took each form.
Question 7: A CBT score report for a certification exam includes a 'score validity period.' What does this communicate?
- The dates the test center is open for testing
- The time window during which the score is officially recognized for credentialing purposes (Correct answer)
- The period a candidate must wait before retesting
- The deadline to receive score reports after testing
Correct answer: The time window during which the score is officially recognized for credentialing purposes
A score validity period specifies how long the score is accepted for credentialing decisions, after which the candidate must retest because knowledge in the field may have changed.
What is the primary advantage of using IRT over Classical Test Theory (CTT) for CBT item calibration?