CBT - Computer Based Testing Psychometric Principles and Analysis Questions and Answers — Questions and Answers
Question 1: A test developer for a new certification exam calculates an item difficulty index (p-value) of 0.92 for a multiple-choice question. What does this value indicate about the item?
- The item is very difficult, as only 8% of test-takers answered it correctly.
- The item has high discrimination, effectively differentiating between high and low performers.
- The item is very easy, as 92% of test-takers answered it correctly. (Correct answer)
- The item may be flawed because the p-value is outside the acceptable range of -1.0 to +1.0.
Correct answer: The item is very easy, as 92% of test-takers answered it correctly.
The item difficulty index (p-value) represents the proportion of test-takers who answered the item correctly. A value of 0.92 means that 92% of the sample answered correctly, indicating that the item is very easy. While easy items can be appropriate for certain goals, this one would likely have low discrimination because most candidates, regardless of ability, answered it correctly.
Question 2: In the context of Computerized Adaptive Testing (CAT), which psychometric theory is most essential for the item selection algorithm?
- Generalizability Theory
- Classical Test Theory (CTT)
- Item Response Theory (IRT) (Correct answer)
- Factor Analysis Theory
Correct answer: Item Response Theory (IRT)
Item Response Theory (IRT) is the foundational psychometric model for CAT. IRT places both test-taker ability and item difficulty on the same scale, allowing the algorithm to select items that are optimally informative for a candidate's estimated ability level. This enables the test to be tailored in real-time.
Question 3: Which of the following metrics provides an estimate of the consistency of a test's results if a person were to take the same test multiple times?
- Item Discrimination Index
- Standard Error of Measurement (SEM) (Correct answer)
- Content Validity Ratio
- Cronbach's Alpha
Correct answer: Standard Error of Measurement (SEM)
The Standard Error of Measurement (SEM) quantifies the amount of error in an individual's test score. It is used to create a confidence interval around the observed score, indicating the likely range where the individual's 'true score' lies. A smaller SEM indicates higher reliability and more confidence that the observed score is close to the true score upon retesting.
Question 4: A psychometrician is analyzing a new test item and finds it has a discrimination index of -0.15. What is the most appropriate action to take?
- Keep the item as it is, as it is functioning well.
- Increase the difficulty of the item to improve discrimination.
- Keep the item, but only for low-stakes practice exams.
- Review the item for flaws, as high-scorers are getting it wrong more often than low-scorers. (Correct answer)
Correct answer: Review the item for flaws, as high-scorers are getting it wrong more often than low-scorers.
A negative discrimination index indicates that lower-scoring test-takers are answering the item correctly more often than higher-scoring test-takers. This is a significant red flag, suggesting the item may be poorly worded, have a confusing key, or be otherwise flawed. The item should be reviewed and likely removed or substantially revised.
Question 5: A certification body is transitioning its exam from a linear, fixed-form test to a Computer Based Test (CBT). They want to ensure that scores are comparable across different versions of the exam administered on different days. This psychometric property is known as:
- Equating (Correct answer)
- Normalization
- Standardization
- Calibration
Correct answer: Equating
Equating is the statistical process used to adjust scores on different forms of a test to make them comparable. This is crucial in CBT environments where different candidates receive different sets of items or test forms, ensuring that a score of, for example, 85 on one form is equivalent in meaning to a score of 85 on another, more or less difficult, form.
Question 6: When comparing Classical Test Theory (CTT) and Item Response Theory (IRT), which statement is most accurate?
- CTT provides item statistics that are independent of the test-taker sample, while IRT statistics are sample-dependent.
- IRT focuses on the total test score as the primary unit of analysis, whereas CTT focuses on individual item responses.
- IRT models the probability of a specific response based on person and item parameters, allowing for more sophisticated analysis than CTT. (Correct answer)
- CTT is a more modern and complex framework, making it better suited for computer adaptive testing than IRT.
Correct answer: IRT models the probability of a specific response based on person and item parameters, allowing for more sophisticated analysis than CTT.
A key advantage of Item Response Theory (IRT) over Classical Test Theory (CTT) is its focus on the item level. IRT uses mathematical models to describe the relationship between a test-taker's latent trait (e.g., ability) and their probability of answering an item correctly. This item-level focus provides more detailed information and is what enables applications like computer adaptive testing. CTT, in contrast, focuses on the properties of the total test score.
A test developer for a new certification exam calculates an item difficulty index (p-value) of 0.92 for a multiple-choice question.
What does this value indicate about the item?