CBT - Computer Based Testing Scoring Models and Reporting Questions and Answers — Questions and Answers
Question 1: A certification body administers multiple forms of a CBT exam throughout the year. To ensure that a score of 450 on Form A represents the same level of proficiency as a score of 450 on the more difficult Form B, which of the following psychometric procedures is essential?
- Standard setting
- Item analysis
- Equating (Correct answer)
- Score validation
Correct answer: Equating
Equating is the statistical process used to adjust scores on different forms of a test to account for variations in difficulty. This ensures that scores are comparable and have the same meaning regardless of which form a candidate takes. Standard setting determines the passing score, item analysis evaluates individual question performance, and score validation is a broader process of ensuring scores are used appropriately.
Question 2: A candidate's score report for a certification exam indicates their scaled score is 750, with a passing score of 700. The report also shows a Standard Error of Measurement (SEM) of 25. What does the SEM indicate?
- The candidate's score is 25 points above the average score of all test-takers.
- The range within which the candidate's 'true score' likely falls. (Correct answer)
- The percentage of questions the candidate answered correctly.
- The percentile rank of the candidate compared to a reference group.
Correct answer: The range within which the candidate's 'true score' likely falls.
The Standard Error of Measurement (SEM) quantifies the amount of error in a test score. It is used to create a confidence interval, or a range of scores, that likely contains the candidate's true score (the score they would get if there were no measurement error). It does not directly relate to the average score, percentage correct, or percentile rank.
Question 3: Which of the following BEST describes the primary advantage of using scaled scores instead of raw scores for reporting results of a high-stakes CBT program with multiple test forms?
- Scaled scores are easier for candidates to understand than raw percentages.
- They allow for fair comparison of candidate performance across different, potentially unequally difficult, test forms. (Correct answer)
- They directly represent the number of questions answered correctly, providing transparent feedback.
- Scaled scores are required for conducting item-level discrimination analysis.
Correct answer: They allow for fair comparison of candidate performance across different, potentially unequally difficult, test forms.
The main purpose of scaling scores is to enable fair and direct comparisons of performance even when candidates take different versions (forms) of an exam that may have minor variations in difficulty. A raw score (number correct) on an easier form is not equivalent to the same raw score on a harder form. Scaling adjusts for this, ensuring that a specific scaled score signifies the same level of proficiency regardless of the form taken.
Question 4: A panel of Subject Matter Experts (SMEs) is convened to determine the passing score for a new certification exam. Each SME independently reviews every test item and estimates the probability that a 'minimally competent candidate' would answer it correctly. The average of these probabilities across all items and all SMEs is then calculated to set the initial pass mark. This process is an application of the:
- Hofstee method
- Contrasting groups method
- Angoff method (Correct answer)
- Bookmark method
Correct answer: Angoff method
This scenario precisely describes the Angoff method, a widely used, item-centered standard-setting procedure. It relies on SME judgments about item difficulty for a hypothetical minimally competent candidate to establish a defensible, criterion-referenced cut score.
Question 5: When designing a score report for candidates who did not pass a CBT exam, which component is MOST valuable for guiding their future study efforts?
- The candidate's overall scaled score and the required passing score.
- A list of the specific questions the candidate answered incorrectly.
- Diagnostic feedback showing performance by major content domains or sections. (Correct answer)
- The candidate's percentile rank compared to all other test-takers.
Correct answer: Diagnostic feedback showing performance by major content domains or sections.
While the overall score indicates the result, diagnostic feedback broken down by content areas (e.g., domains, objectives) provides actionable information. It helps candidates identify their specific areas of weakness so they can focus their remediation efforts effectively. Providing specific incorrect questions is often avoided to protect item security, and percentile rank is less useful for targeted study.
Question 6: A testing organization needs to provide its board of directors with a high-level overview of exam performance trends from the past year, including pass rates by demographic group and geographic region. To do this, individual candidate score data must be compiled and summarized. This process is known as:
- Score scaling
- Data aggregation (Correct answer)
- Standard setting
- Psychometric modeling
Correct answer: Data aggregation
Data aggregation is the process of collecting and combining data from multiple sources to present it in a summary format. In this scenario, individual candidate results are being grouped and summarized to show trends and patterns for stakeholders, which is a classic use case for data aggregation.
A certification body administers multiple forms of a CBT exam throughout the year.
To ensure that a score of 450 on Form A represents the same level of proficiency as a score of 450 on the more difficult Form B, which of the following psychometric procedures is essential?