CBT Scoring Models and Reporting 4 — Questions and Answers
Question 1: What distinguishes horizontal equating from vertical scaling in CBT?
- Horizontal equating compares scores across grade levels; vertical scaling does not
- Horizontal equating links forms of similar difficulty at the same level; vertical scaling links forms across developmental levels (Correct answer)
- Horizontal equating is used only for adaptive tests
- Vertical scaling requires larger sample sizes than horizontal equating
Correct answer: Horizontal equating links forms of similar difficulty at the same level; vertical scaling links forms across developmental levels
Horizontal equating connects parallel forms measuring the same construct at the same level, while vertical scaling links forms across different ability or grade levels.
Question 2: A score report includes a 'confidence band' around a candidate's score. What does this band communicate?
- The range of scores considered passing
- The interval within which the true score likely falls (Correct answer)
- The scores of other candidates who took the same form
- The maximum possible score on the exam
Correct answer: The interval within which the true score likely falls
A confidence band (±1 or ±2 SEM) shows the range where the candidate's true score most likely lies, acknowledging inherent measurement uncertainty.
Question 3: Which reporting practice best supports fairness by ensuring no single demographic group's results are systematically over- or under-reported?
- Reporting only aggregate total scores
- Disaggregating score data by subgroups (Correct answer)
- Using only norm-referenced scores
- Suppressing scores below a certain N-count
Correct answer: Disaggregating score data by subgroups
Disaggregating scores by demographic subgroups allows stakeholders to identify and address performance gaps that may signal bias or inequity.
Question 4: In CBT score reporting, what is 'scale drift' and why is it a concern?
- When scores increase each testing cycle due to coaching
- When the score scale gradually shifts over time, making comparisons across years invalid (Correct answer)
- When candidates score higher on digital than paper tests
- When different testing centers report different passing rates
Correct answer: When the score scale gradually shifts over time, making comparisons across years invalid
Scale drift occurs when equating errors accumulate across testing cycles, causing scores from different years to no longer be truly comparable.
Question 5: A CBT uses a 'mastery model' scoring approach. Which outcome does this model produce?
- A continuous scaled score from 200–800
- A pass/fail determination based on meeting a competency threshold (Correct answer)
- A percentile rank compared to a norm group
- A weighted score adjusted for item difficulty
Correct answer: A pass/fail determination based on meeting a competency threshold
Mastery models classify candidates as either having demonstrated sufficient competency (pass) or not (fail), rather than providing a continuous score.
Question 6: The Contrasting Groups standard-setting method determines cut scores by:
- Having experts rate each item for borderline candidates
- Comparing score distributions of known masters and non-masters (Correct answer)
- Ordering items by difficulty and placing a bookmark
- Averaging panelist ratings across multiple rounds
Correct answer: Comparing score distributions of known masters and non-masters
The Contrasting Groups method finds the cut score where the score distributions of clearly competent and clearly incompetent groups intersect.
Question 7: When a CBT score report displays 'diagnostic feedback,' this typically includes:
- The exact items the candidate answered incorrectly
- Performance summaries by content domain or objective (Correct answer)
- The names of other candidates and their scores
- The adaptive algorithm's item selection pathway
Correct answer: Performance summaries by content domain or objective
Diagnostic feedback aggregates performance by content domain or objective, showing relative strengths and weaknesses without revealing specific item content.
What distinguishes horizontal equating from vertical scaling in CBT?