← All GED Flashcard Decks

Assessment & Evaluation Methods Flashcards

6 cards from real GED practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Assessment & Evaluation Methods flashcards as text
  1. A GED teacher notices that students who perform well on weekly formative quizzes consistently underperform on the official GED Ready practice test. Which psychometric concept best explains this discrepancy?

    Answer: Construct underrepresentation in the formative quizzes

    Construct underrepresentation occurs when an assessment fails to capture the full breadth of the target construct. If the teacher's formative quizzes measure only a narrow slice of the GED content domain, students can master that slice without developing the broader competencies the GED Ready test measures, producing a systematic gap between quiz performance and practice test performance.

  2. When interpreting a student's GED Mathematical Reasoning score using a standard error of measurement (SEM) of 4 points, a score of 150 should be reported as:

    Answer: A true score likely falling between 146 and 154

    The conventional practice is to construct a confidence interval of ±1 SEM around the obtained score to estimate the range likely containing the true score. With SEM = 4, the interval is 150 ± 4, or 146–154. Using ±2 SEM (142–158) represents a 95% confidence interval and is not the standard first-level reporting convention.

  3. A GED preparation program wants to evaluate whether its instructional intervention caused score gains, not merely maturation or regression to the mean. Which design most rigorously supports a causal claim?

    Answer: Randomized control trial with waitlist control group

    A randomized control trial (RCT) with random assignment to treatment or a waitlist control group is the gold standard for establishing causality because it controls for selection bias, maturation, regression to the mean, and historical threats simultaneously. Pre/post single-cohort designs cannot rule out maturation; dropout comparisons introduce severe selection bias; interrupted time-series can suggest but rarely confirm causality at the individual-program level.

  4. A GED teacher uses a rubric to score extended-response answers and finds that two raters assign scores that correlate at r = 0.62. The program director argues this level of inter-rater reliability is acceptable. Which response most accurately challenges that claim?

    Answer: A correlation of 0.62 indicates raters agree on rank order but share only about 38% common variance, which is typically insufficient for high-stakes scoring decisions

    The coefficient of determination (r²) for r = 0.62 is approximately 0.38, meaning raters share only 38% common variance — they disagree substantially more than they agree. For high-stakes decisions like GED scoring, reliability coefficients are generally expected to exceed 0.80–0.85. The claim that 0.62 is acceptable understates the degree of rater inconsistency.

  5. On a GED Science constructed-response item scored 0–3, a teacher's class shows a bimodal distribution with most students scoring either 0 or 3 and very few scoring 1 or 2. This pattern most strongly suggests:

    Answer: A rubric with poorly defined middle-score descriptors, reducing scoring validity

    When raters cluster at score extremes and avoid middle scores, the most common cause is vague or ambiguous descriptors for the intermediate score points (1 and 2). Scorers default to 'clearly right' or 'clearly wrong' because the rubric does not give them enough guidance to differentiate partial credit. This is a rubric design validity problem, not primarily an item difficulty or discrimination issue.

  6. A GED program disaggregates pass rates by racial/ethnic subgroup and finds a 15-percentage-point gap. Before concluding the gap reflects differential instruction quality, which alternative validity threat must be ruled out first?

    Answer: Measurement non-invariance across subgroups, meaning the test may not measure the same construct equally across groups

    Measurement non-invariance (also called differential test functioning at the test level) is the foundational validity threat: if the GED does not measure the same latent construct with the same precision across racial/ethnic groups, then score comparisons across groups are not valid and the gap cannot be meaningfully attributed to instruction. Small-n sampling error is also a concern, but it is secondary — a gap could be statistically stable yet still reflect measurement bias rather than instructional differences.