Psychometric Principles and Analysis Flashcards
6 cards from real CBT practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Psychometric Principles and Analysis flashcards as text
A test developer for a new certification exam calculates an item difficulty index (p-value) of 0.92 for a multiple-choice question. What does this value indicate about the item?
Answer: The item is very easy, as 92% of test-takers answered it correctly.
The item difficulty index (p-value) represents the proportion of test-takers who answered the item correctly. A value of 0.92 means that 92% of the sample answered correctly, indicating that the item is very easy. While easy items can be appropriate for certain goals, this one would likely have low discrimination because most candidates, regardless of ability, answered it correctly.
In the context of Computerized Adaptive Testing (CAT), which psychometric theory is most essential for the item selection algorithm?
Answer: Item Response Theory (IRT)
Item Response Theory (IRT) is the foundational psychometric model for CAT. IRT places both test-taker ability and item difficulty on the same scale, allowing the algorithm to select items that are optimally informative for a candidate's estimated ability level. This enables the test to be tailored in real-time.
Which of the following metrics provides an estimate of the consistency of a test's results if a person were to take the same test multiple times?
Answer: Standard Error of Measurement (SEM)
The Standard Error of Measurement (SEM) quantifies the amount of error in an individual's test score. It is used to create a confidence interval around the observed score, indicating the likely range where the individual's 'true score' lies. A smaller SEM indicates higher reliability and more confidence that the observed score is close to the true score upon retesting.
A psychometrician is analyzing a new test item and finds it has a discrimination index of -0.15. What is the most appropriate action to take?
Answer: Review the item for flaws, as high-scorers are getting it wrong more often than low-scorers.
A negative discrimination index indicates that lower-scoring test-takers are answering the item correctly more often than higher-scoring test-takers. This is a significant red flag, suggesting the item may be poorly worded, have a confusing key, or be otherwise flawed. The item should be reviewed and likely removed or substantially revised.
A certification body is transitioning its exam from a linear, fixed-form test to a Computer Based Test (CBT). They want to ensure that scores are comparable across different versions of the exam administered on different days. This psychometric property is known as:
Answer: Equating
Equating is the statistical process used to adjust scores on different forms of a test to make them comparable. This is crucial in CBT environments where different candidates receive different sets of items or test forms, ensuring that a score of, for example, 85 on one form is equivalent in meaning to a score of 85 on another, more or less difficult, form.
When comparing Classical Test Theory (CTT) and Item Response Theory (IRT), which statement is most accurate?
Answer: IRT models the probability of a specific response based on person and item parameters, allowing for more sophisticated analysis than CTT.
A key advantage of Item Response Theory (IRT) over Classical Test Theory (CTT) is its focus on the item level. IRT uses mathematical models to describe the relationship between a test-taker's latent trait (e.g., ability) and their probability of answering an item correctly. This item-level focus provides more detailed information and is what enables applications like computer adaptive testing. CTT, in contrast, focuses on the properties of the total test score.