M-STEP Item Analysis 3 — Questions and Answers
Question 1: When comparing item analysis results for ELA and math items on the M-STEP, which statement about p-values is accurate?
- P-values should always be identical across content areas
- Acceptable p-value ranges may differ depending on the purpose and content of the test (Correct answer)
- A p-value above 0.50 always indicates a flawed item
- Math items should have higher p-values than ELA items
Correct answer: Acceptable p-value ranges may differ depending on the purpose and content of the test
Acceptable difficulty ranges vary by test purpose; criterion-referenced assessments may tolerate different p-value ranges than norm-referenced tests.
Question 2: Which scenario BEST describes an item exhibiting uniform DIF favoring one student group?
- The item is harder for one group at all ability levels compared to another group (Correct answer)
- The item is harder for one group only at high ability levels
- The item has a low discrimination index for both groups
- The item has different p-values only when the groups are combined
Correct answer: The item is harder for one group at all ability levels compared to another group
Uniform DIF occurs when one group consistently performs better or worse than the other across all ability levels, not just at certain points.
Question 3: A test developer sees that an item's discrimination index is 0.00. What is the most likely explanation?
- High scorers and low scorers answered the item correctly at the same rate (Correct answer)
- The item was answered incorrectly by all students
- The item has too many answer choices
- The item was skipped by most students
Correct answer: High scorers and low scorers answered the item correctly at the same rate
A discrimination index of 0.00 means there is no difference in the proportion of high scorers versus low scorers who answered correctly.
Question 4: Why might a constructed-response item on the M-STEP require inter-rater reliability analysis in addition to standard item analysis?
- Constructed-response items always have higher p-values than multiple-choice items
- Scores depend on human judgment, introducing potential scoring inconsistency (Correct answer)
- Constructed-response items cannot be analyzed using classical test theory
- Inter-rater reliability replaces the need for a discrimination index
Correct answer: Scores depend on human judgment, introducing potential scoring inconsistency
Because human raters score constructed-response items, inter-rater reliability measures how consistently different raters apply the scoring rubric.
Question 5: Which statement accurately describes the relationship between item difficulty and test reliability?
- Items with p-values near 0.50 tend to maximize score variance and support higher reliability (Correct answer)
- Easier items always produce more reliable tests
- Reliability is unaffected by the difficulty of individual items
- Harder items always improve test reliability
Correct answer: Items with p-values near 0.50 tend to maximize score variance and support higher reliability
Items near p = 0.50 produce maximum variance in student responses, which increases the test's overall score variance and supports higher reliability coefficients.
Question 6: An item analysis report shows that 40% of students in the upper group and 38% in the lower group answered Item 12 correctly. What does this suggest?
- Item 12 is highly discriminating
- Item 12 should be used as the anchor item for scaling
- Item 12 is a poor discriminator and may need revision (Correct answer)
- Item 12 has a p-value that is too low
Correct answer: Item 12 is a poor discriminator and may need revision
When the upper and lower groups perform nearly identically on an item, the discrimination index is close to zero, indicating the item fails to differentiate ability levels.
Question 7: How does item response theory (IRT) improve on classical test theory (CTT) when analyzing M-STEP items?
- IRT eliminates the need for large sample sizes
- IRT provides item parameter estimates that are sample-independent (Correct answer)
- IRT does not require scoring rubrics for open-ended items
- IRT uses only p-values to evaluate item quality
Correct answer: IRT provides item parameter estimates that are sample-independent
A key advantage of IRT is that item parameters (difficulty, discrimination) are theoretically invariant across different samples of test-takers.
When comparing item analysis results for ELA and math items on the M-STEP, which statement about p-values is accurate?