M-STEP Item Analysis 4 — Questions and Answers
Question 1: During M-STEP item review, an item is flagged for C-DIF (non-uniform DIF). What does this mean?
- The item favors one group consistently at all ability levels
- The group performance gap changes direction or magnitude depending on ability level (Correct answer)
- The item's p-value differs between content areas
- The item was answered correctly by fewer than 30% of students
Correct answer: The group performance gap changes direction or magnitude depending on ability level
Non-uniform DIF means the interaction between group membership and ability level is not constant — the relative advantage switches or changes across the ability scale.
Question 2: A set of M-STEP items has an average point-biserial correlation of 0.45. What does this suggest about the item set?
- The items are too easy for the intended grade level
- The items are highly reliable but have low content validity
- The items generally discriminate well between higher and lower achievers (Correct answer)
- The items should be removed from the operational test immediately
Correct answer: The items generally discriminate well between higher and lower achievers
Point-biserial correlations averaging around 0.45 indicate items that effectively differentiate students with higher and lower overall test performance.
Question 3: Which of the following is a limitation of using only classical test theory statistics for M-STEP item analysis?
- CTT cannot compute p-values for constructed-response items
- Item difficulty and discrimination statistics depend on the specific sample of students tested (Correct answer)
- CTT overestimates the number of functioning distractors
- CTT requires a minimum of 10,000 students for valid results
Correct answer: Item difficulty and discrimination statistics depend on the specific sample of students tested
In CTT, item statistics such as p-values and discrimination indices are sample-dependent, meaning they may differ across different groups of test-takers.
Question 4: What is the purpose of flagging items for sensitivity review as part of item analysis?
- To identify items that are too difficult for below-grade-level students
- To ensure items do not contain content that could offend or disadvantage specific demographic groups (Correct answer)
- To check that items align with state curriculum standards
- To verify that items have at least four answer choices
Correct answer: To ensure items do not contain content that could offend or disadvantage specific demographic groups
Sensitivity review identifies content that might be culturally biased, offensive, or unfair to specific groups of students, independent of statistical bias measures.
Question 5: An educator analyzing M-STEP data notices that Item 7 has a p-value of 0.30 and a discrimination index of 0.40. Which conclusion is BEST supported?
- Item 7 should be eliminated because it is too difficult
- Item 7 is hard but effectively distinguishes high from low performers (Correct answer)
- Item 7 has weak discrimination and poor content validity
- Item 7 is appropriate only for gifted students
Correct answer: Item 7 is hard but effectively distinguishes high from low performers
A discrimination index of 0.40 is quite strong; even with a low p-value, this item meaningfully separates students by achievement level.
Question 6: What does the term 'item parameter drift' mean in the context of longitudinal M-STEP item analysis?
- An item becomes harder over time due to curriculum changes
- An item's statistical parameters shift significantly from one testing year to the next (Correct answer)
- An item's wording changes between operational forms
- An item begins to discriminate against students with disabilities over time
Correct answer: An item's statistical parameters shift significantly from one testing year to the next
Item parameter drift refers to significant changes in an item's difficulty or discrimination parameters across administrations, which can threaten score comparability.
Question 7: Which action should a test developer take when an item shows large DIF favoring one subgroup but is also highly discriminating and content-valid?
- Automatically remove the item from the test bank
- Conduct a content review to determine whether the DIF reflects real ability differences or construct-irrelevant bias (Correct answer)
- Increase the item's weight in the total score
- Replace the item with a technology-enhanced item
Correct answer: Conduct a content review to determine whether the DIF reflects real ability differences or construct-irrelevant bias
Large DIF does not automatically indicate bias; a content review helps determine whether the group difference reflects genuine differences in the measured construct or unfair item characteristics.
During M-STEP item review, an item is flagged for C-DIF (non-uniform DIF).
What does this mean?