← All GED Flashcard Decks

Assessment & Evaluation Methods Flashcards

6 cards from real GED practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Assessment & Evaluation Methods flashcards as text
  1. A GED teacher notices that students from one demographic group consistently score lower on a constructed-response writing prompt, even though classroom observations show equivalent skill levels across groups. Which psychometric concern does this MOST likely indicate?

    Answer: Differential item functioning (DIF), suggesting the item may be measuring construct-irrelevant factors tied to group membership

    Differential item functioning (DIF) occurs when examinees of equal ability but different group membership respond differently to an item, suggesting the item captures something beyond the intended construct (e.g., cultural familiarity with a topic). This is distinct from overall score gaps, which may reflect real skill differences. DIF analysis is the appropriate diagnostic tool here, not reliability or content validity concerns, since the teacher's observation controls for actual skill level.

  2. A GED program coordinator wants to evaluate whether the instructional program causes gains in student achievement, not just whether students who attend more happen to score higher. Which evaluation design BEST addresses this internal validity threat?

    Answer: A pre-test/post-test design with a matched comparison group of non-enrollees

    The threat described is selection bias — students who attend more may already be more motivated or able. A matched comparison group (quasi-experimental design) controls for pre-existing differences by comparing similar non-enrollees, isolating the effect of the program itself. Simple correlations conflate attendance with ability; normed tests and satisfaction surveys don't address causality at all.

  3. When interpreting a student's score on the GED Ready™ practice test, a teacher sees a score in the 'Likely to Pass' range with a wide confidence interval. What is the MOST technically accurate interpretation of the confidence interval in this context?

    Answer: If the student took many equivalent forms of the test, approximately 90% of the resulting scores would fall within that interval, reflecting measurement error around the observed score

    A confidence interval around a test score is a frequentist construct: it reflects that if we repeatedly sampled equivalent test forms, the interval would contain the true score approximately 90% of the time. It is NOT a probability statement about this one student's true score (a common misconception), nor is it a reliability coefficient or a norm-referenced range. The interval is built from the standard error of measurement (SEM).

  4. A GED teacher designs a portfolio assessment where students select their best work samples as evidence of competency. A colleague argues this approach introduces a specific validity threat. Which threat is MOST applicable, and why?

    Answer: Construct underrepresentation, because student self-selection may systematically exclude evidence of weaker skill areas, producing an incomplete picture of actual competency

    When students self-select portfolio entries, they naturally choose their strongest work, which may omit evidence from skill domains they find difficult. This creates construct underrepresentation — the portfolio no longer adequately samples the full competency domain. Consequential invalidity concerns effects of score use, not the score itself; halo effect is a scorer bias, not a sampling problem; criterion contamination means the criterion is influenced by the predictor, which isn't the issue here.

  5. A GED instructor uses a rubric with four analytic dimensions (Ideas, Organization, Voice, Conventions) and finds that scores on all four dimensions correlate at r = .95 with each other across all student papers. What does this MOST likely indicate about the rubric's construct validity?

    Answer: The dimensions may lack discriminant validity — they are likely measuring the same underlying factor rather than four distinct constructs

    When analytically scored dimensions correlate at r = .95, they are essentially interchangeable — a student's score on Ideas predicts their Conventions score almost perfectly. This indicates the rubric is not capturing four distinct constructs but rather one general writing quality factor. This is a discriminant validity failure: the dimensions should correlate moderately (showing they measure related but distinct facets) but not this highly. High intercorrelation is not evidence of inter-rater reliability (which requires comparing two scorers on the same dimension) or predictive validity.

  6. A program evaluator conducting a formative evaluation of a GED preparation curriculum receives interim assessment data showing that students score well on recall items but poorly on transfer items that require applying concepts to novel contexts. Which curricular inference is MOST defensible from this pattern?

    Answer: Instruction may be emphasizing surface-level encoding strategies without sufficient elaborative practice that promotes flexible knowledge transfer

    The gap between recall and transfer performance is a classic signal of instruction focused on rote encoding rather than deep, elaborative processing. Transfer requires flexible mental models built through varied practice, retrieval in novel contexts, and interleaving — not just repeated exposure to the same content format. Removing transfer items undermines valid measurement of the GED's emphasis on applied reasoning. More time-on-task with the same approach would not close the gap if the approach itself doesn't support transfer. 'Floor effect' refers to scores clustering at the bottom, which contradicts the premise that recall scores are high.