ABLE Technical Foundations 3 — Questions and Answers
Question 1: Which item response theory (IRT) model uses only item difficulty as its parameter?
- 2PL model
- 3PL model
- Rasch (1PL) model (Correct answer)
- Generalizability model
Correct answer: Rasch (1PL) model
The Rasch model (one-parameter logistic model) estimates examinee ability based solely on item difficulty, assuming equal discrimination across items.
Question 2: A normal distribution of test scores has a mean of 70 and a standard deviation of 10. Approximately what percentage of scores fall between 60 and 80?
- 95%
- 99%
- 68% (Correct answer)
- 50%
Correct answer: 68%
In a normal distribution, approximately 68% of scores fall within one standard deviation above and below the mean.
Question 3: Which type of language test requires examinees to produce unrestricted written or spoken responses scored by raters?
- Selected-response test
- Constructed-response test (Correct answer)
- Matching test
- True/false test
Correct answer: Constructed-response test
Constructed-response tests require examinees to produce their own responses, which are then evaluated by trained raters using scoring criteria.
Question 4: What is the primary purpose of anchor items in equating across multiple test forms?
- To increase test security
- To link scores from different forms to a common scale (Correct answer)
- To measure test-taker motivation
- To reduce total test length
Correct answer: To link scores from different forms to a common scale
Anchor items appear on multiple test forms and are used to statistically link and equate the score scales of those forms.
Question 5: Inter-rater reliability is BEST measured by which statistic when ratings are on an ordinal scale?
- Pearson r
- Cohen's kappa (Correct answer)
- Cronbach's alpha
- Point-biserial correlation
Correct answer: Cohen's kappa
Cohen's kappa measures agreement between two raters on categorical or ordinal data while correcting for agreement due to chance.
Question 6: Which principle of language assessment states that a test should reflect real-world language use tasks?
- Reliability
- Authenticity (Correct answer)
- Interactiveness
- Impact
Correct answer: Authenticity
Authenticity in language testing means that test tasks should closely resemble the target language use situations examinees encounter in real life.
Question 7: What does the standard error of measurement (SEM) indicate about a test score?
- The average score across all test forms
- The expected variability in an individual's scores across repeated administrations (Correct answer)
- The correlation between two parallel forms
- The percentage of items answered correctly
Correct answer: The expected variability in an individual's scores across repeated administrations
The SEM quantifies the amount of error in an individual score, indicating how much that score might fluctuate if the person were tested again under the same conditions.
Which item response theory (IRT) model uses only item difficulty as its parameter?