Evaluation Format Flashcards
7 cards from real GAT practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Evaluation Format flashcards as text
During GAT administration, 'field-test items' are included in the exam but do not count toward the score. Their purpose is to:
Answer: Collect data to evaluate whether items are suitable for future scored use
Field-test items are piloted to gather statistical information before being used as scored operational items.
A GAT section uses 'polytomous scoring,' meaning:
Answer: Items can earn partial credit with multiple score levels
Polytomous scoring assigns multiple score levels to an item, allowing for partial credit rather than all-or-nothing scoring.
The GAT evaluation is described as 'criterion-referenced' for certain sections. This means student performance is compared to:
Answer: A pre-defined standard or set of learning outcomes
Criterion-referenced tests measure performance against defined content standards, not against other test-takers.
Which evaluation format feature ensures that every test-taker, regardless of disability, has fair access to the GAT?
Answer: Standardized accommodations
Standardized accommodations such as extended time or alternative formats provide equitable access for test-takers with disabilities.
A GAT score is said to have 'content validity' when:
Answer: The test items adequately represent the full range of skills being assessed
Content validity ensures that test items comprehensively cover the domain or subject matter the test is designed to assess.
On the GAT, when test directions state 'omitted answers will not be penalized,' test-takers should understand that:
Answer: No score deduction is applied for unanswered items
Without a guessing penalty, skipping items simply means no points earned or lost for those questions.
Which of the following best describes 'inter-rater reliability' in GAT constructed-response scoring?
Answer: The degree of agreement between two or more raters scoring the same response
Inter-rater reliability measures how consistently different raters assign the same score to the same response.