CBT Psychometric Principles and Analysis 4 — Questions and Answers
Question 1: A psychometrician calculates the standard error of measurement (SEM) for a test with SD = 15 and reliability = 0.91. Approximately what is the SEM?
- 1.35
- 4.50 (Correct answer)
- 13.65
- 6.75
Correct answer: 4.50
SEM = SD × √(1 − reliability) = 15 × √(0.09) = 15 × 0.30 = 4.50.
Question 2: Which of the following is a key assumption of parallel test forms in classical test theory?
- Both forms must have identical items
- Both forms have equal true scores and equal error variances for all examinees (Correct answer)
- Both forms must be administered to different samples
- Item response functions must be identical across forms
Correct answer: Both forms have equal true scores and equal error variances for all examinees
Parallel forms require equal true scores and equal error variances for every examinee, ensuring the forms are interchangeable measures of the same construct.
Question 3: In computerized adaptive testing (CAT), which method is most commonly used to select the next item after each response?
- Random item selection from the item bank
- Select the item with maximum information at the current ability estimate (Correct answer)
- Select the easiest unadministered item
- Select the item with the highest content validity ratio
Correct answer: Select the item with maximum information at the current ability estimate
CAT algorithms typically select the item that provides maximum Fisher information at the current theta estimate, efficiently narrowing the measurement error.
Question 4: A test blueprint specifies that 30% of items must cover 'Application' and 20% must cover 'Analysis' per Bloom's Taxonomy. This document primarily supports which type of validity evidence?
- Criterion-related validity
- Content validity (Correct answer)
- Convergent validity
- Consequential validity
Correct answer: Content validity
A test blueprint or table of specifications documents systematic domain coverage, providing content validity evidence that the test adequately samples the intended construct.
Question 5: When comparing two tests measuring the same construct, a correlation of r = 0.78 is found. This supports which type of validity evidence?
- Discriminant validity
- Convergent validity (Correct answer)
- Incremental validity
- Face validity
Correct answer: Convergent validity
Convergent validity evidence is demonstrated when a test correlates highly with other measures of the same or theoretically related constructs.
Question 6: The Rasch model is considered a special case of which broader IRT model?
- The two-parameter logistic model (2PL)
- The three-parameter logistic model (3PL)
- The one-parameter logistic model (1PL), which is equivalent to Rasch (Correct answer)
- The graded response model (GRM)
Correct answer: The one-parameter logistic model (1PL), which is equivalent to Rasch
The Rasch model is mathematically equivalent to the 1PL model, constraining all item discriminations to be equal and not including a guessing parameter.
Question 7: A test is described as having high 'ecological validity.' What does this mean?
- The test was normed on a large, nationally representative sample
- Test performance generalizes well to real-world behaviors and settings (Correct answer)
- The test has minimal environmental scoring bias
- Items were derived from naturalistic observation
Correct answer: Test performance generalizes well to real-world behaviors and settings
Ecological validity refers to how well test scores predict or reflect functioning in real-world, everyday contexts beyond the testing environment.
A psychometrician calculates the standard error of measurement (SEM) for a test with SD = 15 and reliability = 0.91.
Approximately what is the SEM?