M-STEP Item Analysis 5 — Questions and Answers
Question 1: Which piece of item analysis information is MOST useful for evaluating whether a specific wrong answer is serving as an effective distractor?
- The item's IRT b-parameter
- The proportion of students selecting each response option (Correct answer)
- The item's standard error of measurement
- The item's content alignment rating
Correct answer: The proportion of students selecting each response option
Distractor analysis examines what percentage of students chose each option, revealing whether incorrect choices are attracting meaningful numbers of test-takers.
Question 2: A new M-STEP item shows that high-scoring students more frequently chose distractor B than the keyed correct answer. What is the MOST likely explanation?
- Distractor B is too implausible
- The item may be miskeyed or contain a technical flaw (Correct answer)
- High-scoring students are guessing more often
- The item has a very low p-value
Correct answer: The item may be miskeyed or contain a technical flaw
When high scorers disproportionately choose a distractor over the key, it strongly suggests the item is miskeyed or contains ambiguous language that makes the distractor appear correct.
Question 3: How does sample size affect the reliability of item analysis statistics for field-test items on the M-STEP?
- Larger samples produce more stable and trustworthy item statistics (Correct answer)
- Smaller samples always yield higher discrimination indices
- Sample size affects p-values but not discrimination indices
- Item analysis statistics are equally reliable regardless of sample size
Correct answer: Larger samples produce more stable and trustworthy item statistics
Larger sample sizes reduce sampling error, making item statistics like p-values and discrimination indices more stable and reliable for decision-making.
Question 4: Which IRT model parameter directly corresponds to item difficulty in classical test theory?
- The a-parameter (discrimination)
- The b-parameter (difficulty/location) (Correct answer)
- The c-parameter (pseudo-guessing)
- The d-parameter (scaling constant)
Correct answer: The b-parameter (difficulty/location)
In IRT, the b-parameter represents the point on the ability scale where a student has a 50% probability of answering correctly, which is analogous to item difficulty.
Question 5: When analyzing M-STEP writing items scored on a 4-point rubric, which statistic is MOST appropriate for evaluating inter-rater consistency?
- Pearson correlation between scores and item difficulty
- Weighted kappa or percentage exact agreement between raters (Correct answer)
- Point-biserial correlation between item scores and total scores
- Cronbach's alpha across all writing items
Correct answer: Weighted kappa or percentage exact agreement between raters
Weighted kappa and percentage of exact agreement are standard metrics for measuring how consistently two human raters assign the same score to the same response.
Question 6: An item bank review reveals that Item 22 has been on operational M-STEP forms for six consecutive years with a steadily declining p-value. What is the MOST likely cause?
- The item's answer key was changed in year three
- Classroom instruction has shifted away from the content the item assesses
- The item is experiencing item exposure effects as students become familiar with it (Correct answer)
- The item's discrimination index has increased over time
Correct answer: The item is experiencing item exposure effects as students become familiar with it
Overexposed items become known to students through test preparation materials and student sharing, inflating scores initially; a declining p-value over years of use can also reflect curriculum misalignment, but exposure effects are a primary concern.
Question 7: Which of the following BEST describes the purpose of anchor items in M-STEP equating?
- Anchor items set the average difficulty for each test form
- Anchor items appear on multiple forms and link them to a common scale (Correct answer)
- Anchor items are the easiest items included to boost student confidence
- Anchor items are used exclusively for DIF analysis
Correct answer: Anchor items appear on multiple forms and link them to a common scale
Anchor (or common) items appear on multiple test forms and allow psychometricians to place scores from different forms onto the same scale through statistical equating.
Which piece of item analysis information is MOST useful for evaluating whether a specific wrong answer is serving as an effective distractor?