Test Validity and Reliability Flashcards
7 cards from real TAPAS practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Test Validity and Reliability flashcards as text
The standard error of measurement (SEM) for a TAPAS scale is most useful for:
Answer: Constructing confidence intervals around an individual's observed score
The SEM estimates the amount of error in an individual score and is used to build a confidence interval showing the probable range of a person's true score.
A TAPAS validity study finds that the Dominance scale correlates r=0.60 with peer-rated leadership but only r=0.10 with a measure of clerical speed. This pattern of results supports:
Answer: Convergent and discriminant validity
High correlation with theoretically related criteria (leadership) and low correlation with unrelated criteria (clerical speed) together constitute the convergent-discriminant validity pattern described by the multitrait-multimethod approach.
In Item Response Theory (IRT), the parameter that describes how well a TAPAS item distinguishes between examinees at different trait levels is the:
Answer: Discrimination parameter (a)
The discrimination parameter (a) in IRT reflects the slope of the item characteristic curve; higher values indicate the item better separates individuals at different levels of the latent trait.
Which of the following would most threaten the content validity of a TAPAS scale designed to measure integrity?
Answer: Including items that assess rule-following but omitting items assessing honesty
Content validity requires that items systematically cover the full domain of the construct; omitting a key facet (honesty) of integrity would leave the scale content-deficient.
A researcher re-administers TAPAS to the same soldiers 6 months later and finds low correlations between Time 1 and Time 2 scores. Which explanation is most consistent with a validity interpretation rather than a reliability interpretation?
Answer: Personality itself changed meaningfully over 6 months of military experience
If personality genuinely changed due to military socialization, low test-retest correlations would reflect real trait change rather than measurement unreliability, which is a validity-relevant finding.
Which of the following is a primary advantage of using an adaptive administration format in TAPAS compared to a fixed-length format?
Answer: Fewer items are needed to achieve equivalent measurement precision
Computerized adaptive testing (CAT) tailors item selection to each examinee's estimated trait level, achieving comparable precision with fewer items than a fixed-length test.
Differential item functioning (DIF) analysis in TAPAS is primarily used to:
Answer: Identify items that function differently for groups matched on the underlying trait
DIF analysis detects items that show different response probabilities for members of different groups (e.g., men vs. women) even when those groups are equal on the latent trait, which signals potential bias.