Assessment & Evaluation Methods Flashcards
6 cards from real GED practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Assessment & Evaluation Methods flashcards as text
A GED teacher administers the same reading comprehension assessment twice to the same cohort, six weeks apart, without intervening instruction. The correlation between the two sets of scores is r = 0.91. Which measurement property does this PRIMARILY establish, and what is its key limitation in this context?
Answer: Test-retest reliability; it cannot confirm the construct being measured is stable rather than the test itself being predictable
A r = 0.91 correlation between two administrations of the same test over time is the definition of test-retest reliability. Its critical limitation is the stability assumption: high correlation may reflect that the construct (e.g., reading ability) is genuinely stable, but it could also mean examinees simply remembered item cues — a threat called 'memory contamination.' It does not establish what the test measures, only that scores are consistent over time.
When interpreting a GED candidate's Extended Response score, an evaluator notices the rubric assigns equal weight to 'Development of Ideas' and 'Language Facility & Conventions.' A candidate scores 3/3 on conventions but 0/3 on development. Which validity threat is MOST directly implicated if the evaluator infers the candidate is 'a competent writer'?
Answer: Construct-irrelevant variance, because conventions performance is inflating the overall writing inference
Construct-irrelevant variance occurs when a test score is influenced by something outside the intended construct. Here, 'competent writer' is the inference, but the candidate's strong conventions score — a peripheral mechanical skill — is driving that interpretation despite zero development of ideas. The evaluator is allowing an irrelevant dimension to elevate the construct inference. Construct underrepresentation would apply if the test simply missed dimensions, but in this case the dimension IS measured; the problem is the evaluator's selective over-reliance on one subscale.
A GED preparation program uses a cut score of 145 on each subject test. A psychometrician applies the Angoff method to recommend a cut score but finds high inter-rater variance among the teacher panelists. Which corrective action BEST addresses this specific methodological weakness?
Answer: Switch to the Bookmark method, which anchors judgments to item response theory difficulty parameters rather than relying solely on panelist intuition
The Angoff method asks panelists to estimate the probability that a minimally competent examinee answers each item correctly — a highly intuitive and often inconsistent judgment process, leading to high inter-rater variance. The Bookmark method mitigates this by ordering items by IRT-derived difficulty and having panelists place a 'bookmark' where minimal competency ends, grounding judgments in empirical difficulty data rather than pure estimation. This reduces reliance on panelist intuition and typically produces greater inter-rater consistency.
A GED teacher compares two groups of students: those who passed the Mathematical Reasoning test on the first attempt versus those who required three attempts. She finds the first-attempt group scores significantly higher on a spatial reasoning pre-test administered before any GED preparation. This finding MOST directly supports which type of validity evidence?
Answer: Known-groups validity evidence, because the pre-test discriminates between groups with theoretically different expected performance
Known-groups validity (a form of criterion-related or construct validity evidence) is established when an instrument differentiates between groups that theory predicts should differ. Here, first-attempt passers and multi-attempt candidates are theoretically different in mathematical preparedness, and the spatial reasoning pre-test successfully separates them — supporting the inference that the pre-test meaningfully measures something relevant to mathematical competence. Convergent validity would require correlation evidence with a similar measure, not group differentiation.
In a GED classroom formative assessment cycle, a teacher uses 'exit tickets' in which students self-rate their confidence on newly taught concepts (1–5 scale). Research consistently shows that lower-skilled students tend to overestimate their competence on such self-assessments. Which assessment design modification MOST directly counteracts this specific bias without eliminating student self-reflection?
Answer: Pair the confidence rating with a single concrete performance item on the same concept, so self-rating can be calibrated against actual demonstrated performance
The Dunning-Kruger effect describes how low-competence individuals systematically overestimate their ability because they lack the metacognitive skill to recognize their own errors. Pairing the confidence rating with an actual performance item (e.g., 'solve this problem AND rate your confidence') creates a calibration signal: the teacher sees both the self-perception AND the evidence of actual competence, allowing targeted feedback that helps students develop metacognitive accuracy over time. This preserves self-reflection while adding an objective anchor to the subjective rating.
A GED testing center reports that examinees who receive extended-time accommodations on the Mathematical Reasoning test score, on average, 8 points higher than their score without accommodation in a pilot study. A validity researcher argues this finding is AMBIGUOUS with respect to accommodation validity. Which interpretation BEST explains the ambiguity?
Answer: The score gain could reflect removal of construct-irrelevant variance for examinees with disabilities, OR it could reflect an unfair advantage that inflates scores beyond what the construct requires, and the data alone cannot distinguish these two explanations
Accommodation validity rests on the 'interaction hypothesis': a valid accommodation should boost scores for examinees with the relevant disability (by removing construct-irrelevant barriers like processing speed) but NOT boost scores for examinees without the disability. If extended time raises scores for all examinees, it may be removing a construct-relevant speed component, thereby changing what the test measures rather than leveling the playing field. The pilot data showing an 8-point gain is ambiguous because it cannot, without a disability-status × condition interaction design, tell us whether the gain reflects valid barrier removal or construct alteration.