KSA Assessment Design & Evaluation Methods 3 — Questions and Answers
Question 1: The standard error of measurement (SEM) is used in assessment to:
- Compare two tests' difficulty levels
- Estimate the range of error around an observed score (Correct answer)
- Calculate the test's internal consistency
- Determine the number of items needed
Correct answer: Estimate the range of error around an observed score
The SEM quantifies measurement error and is used to build a confidence interval around an examinee's observed score.
Question 2: A structured behavioral interview differs from a traditional interview primarily because it:
- Focuses on hypothetical future scenarios only
- Uses standardized questions and a scoring rubric based on past behavior (Correct answer)
- Allows interviewers to probe freely based on candidate responses
- Excludes job-related KSAs from questions
Correct answer: Uses standardized questions and a scoring rubric based on past behavior
Structured behavioral interviews use consistent, pre-determined questions tied to job KSAs and score responses using an anchored rating scale.
Question 3: Which analysis identifies whether a test item functions differently for subgroups of similar ability, potentially indicating bias?
- Item Response Theory (IRT) fit analysis
- Differential Item Functioning (DIF) analysis (Correct answer)
- Classical Test Theory (CTT) difficulty index
- Factor analysis
Correct answer: Differential Item Functioning (DIF) analysis
DIF analysis detects items where examinees of equal ability from different demographic groups have different probabilities of answering correctly, signaling potential bias.
Question 4: A summative evaluation of a training program is conducted:
- Before the program to assess learner readiness
- During the program to guide real-time adjustments
- After the program to judge overall effectiveness (Correct answer)
- At program midpoint to revise materials
Correct answer: After the program to judge overall effectiveness
Summative evaluation assesses the overall impact and outcomes of a completed program, unlike formative evaluation which informs ongoing improvements.
Question 5: In the Angoff method for setting cut scores, SMEs are asked to:
- Rank all examinees by performance
- Estimate the probability that a minimally competent candidate answers each item correctly (Correct answer)
- Calculate the mean item difficulty of the test
- Review test items for content coverage only
Correct answer: Estimate the probability that a minimally competent candidate answers each item correctly
The Angoff method has SMEs judge the probability that a borderline-competent candidate would correctly answer each test item, and these probabilities are summed to establish the cut score.
Question 6: Which of the following best describes the purpose of a Table of Specifications (TOS) in assessment design?
- Documenting the statistical properties of each item
- Mapping test content to cognitive levels and topic areas (Correct answer)
- Recording candidate demographic information
- Tracking the historical performance of an item bank
Correct answer: Mapping test content to cognitive levels and topic areas
A Table of Specifications aligns test items to specific content domains and cognitive levels (e.g., Bloom's Taxonomy) to ensure balanced and valid coverage.
Question 7: An assessment battery administered to candidates yields scores on five separate tests. A composite score is formed by weighting each test. This approach is best described as:
- Multiple hurdles selection
- Compensatory selection (Correct answer)
- Clinical prediction
- Top-down ranking on a single measure
Correct answer: Compensatory selection
A compensatory model combines scores across predictors so that a high score on one measure can offset a low score on another when forming a composite.
The standard error of measurement (SEM) is used in assessment to: