SIOP Psychological Assessment & Measurement Techniques 5 — Questions and Answers
Question 1: The Cleary model of test fairness defines a test as fair when:
- Adverse impact is absent for all subgroups
- The regression lines for predicting criterion performance are identical across subgroups (Correct answer)
- Mean scores are equal across demographic groups
- The reliability coefficients are equal across subgroups
Correct answer: The regression lines for predicting criterion performance are identical across subgroups
The Cleary model defines test fairness as the absence of differential prediction: the same regression line (same slope and intercept) predicts criterion performance for all subgroups.
Question 2: Which statistical method is used to detect differential item functioning (DIF) in large-scale employment tests?
- Principal components analysis
- Mantel-Haenszel procedure (Correct answer)
- Multiple regression with interaction terms
- Confirmatory factor analysis
Correct answer: Mantel-Haenszel procedure
The Mantel-Haenszel procedure compares item response probabilities across groups (e.g., focal vs. reference) matched on total score, making it a standard DIF detection method.
Question 3: A structured behavioral interview achieves higher validity than an unstructured interview primarily because structure:
- Allows interviewers to build rapport with candidates
- Reduces interviewer error and bias by standardizing questions and rating scales (Correct answer)
- Increases the length of the interview
- Focuses on candidates' future intentions rather than past behavior
Correct answer: Reduces interviewer error and bias by standardizing questions and rating scales
Structure reduces idiosyncratic interviewer effects, anchors ratings to observable behaviors, and standardizes the evaluation process, all of which improve reliability and validity.
Question 4: The concept of 'synthetic validity' (job component validity) is most useful when:
- A large sample of incumbents is available for criterion-related validation
- Sample sizes are too small for traditional criterion-related validation in specific jobs (Correct answer)
- Adverse impact must be eliminated from a selection system
- Content validity evidence alone is required by the EEOC
Correct answer: Sample sizes are too small for traditional criterion-related validation in specific jobs
Synthetic validity links validity evidence from multiple job elements across jobs, allowing small organizations to build validity evidence where single-job samples are too small for traditional validation.
Question 5: In a work sample test, which type of validity evidence is PRIMARILY demonstrated?
- Predictive criterion-related validity
- Content validity through job representativeness (Correct answer)
- Construct validity via factor analysis
- Concurrent validity using incumbent scores
Correct answer: Content validity through job representativeness
Work sample tests derive their validity primarily from content: they directly replicate critical job tasks, so content validity evidence from job analysis is the primary justification.
Question 6: The utility formula developed by Brogden, Cronbach, and Gleser estimates the dollar value of a selection system. Which variable has the GREATEST practical impact on utility in most organizational contexts?
- Selection ratio
- Test validity coefficient (r_xy)
- Standard deviation of job performance in dollars (SD_y) (Correct answer)
- Number of applicants tested
Correct answer: Standard deviation of job performance in dollars (SD_y)
SD_y (the standard deviation of job performance expressed in dollars) has the greatest practical impact because utility is directly proportional to it, and it is often the largest and most variable parameter.
Question 7: Which response set involves test-takers consistently choosing extreme response options (e.g., 'strongly agree' or 'strongly disagree') on Likert-scale items regardless of item content?
- Acquiescence bias
- Extreme response style (Correct answer)
- Social desirability bias
- Midpoint responding
Correct answer: Extreme response style
Extreme response style is the tendency to select the most extreme endpoints of rating scales, which introduces systematic measurement error unrelated to the construct being assessed.
The Cleary model of test fairness defines a test as fair when: