TEA Assessment & Evaluation Methods 2 — Questions and Answers
Question 1: A principal notices that students from one demographic group consistently score lower on a standardized test than other groups, despite similar classroom performance. What should the principal investigate first?
- Differential item functioning to check for assessment bias (Correct answer)
- Whether teachers are grading those students more leniently
- The number of absences among students in that group
- Whether the test was administered at a different time of day
Correct answer: Differential item functioning to check for assessment bias
Differential item functioning (DIF) analysis identifies whether specific test items perform differently across demographic groups, revealing potential cultural or linguistic bias.
Question 2: A campus uses a universal screener three times per year to identify students at risk for reading difficulties. This practice best exemplifies which assessment approach?
- Summative evaluation
- Universal screening within an RTI/MTSS framework (Correct answer)
- Criterion-referenced testing
- Norm-referenced benchmarking
Correct answer: Universal screening within an RTI/MTSS framework
Universal screening conducted at regular intervals to identify at-risk students is a core component of Response to Intervention (RTI) and Multi-Tiered System of Supports (MTSS).
Question 3: When reviewing STAAR results, a principal wants to determine how the campus performs relative to other schools with similar student demographics. Which comparison is most appropriate?
- Comparing to the state average for all campuses
- Using a demographically similar campus comparison group (Correct answer)
- Comparing only to the district average
- Reviewing only the campus's own year-over-year trend
Correct answer: Using a demographically similar campus comparison group
Comparing to demographically similar campuses (peer groups) controls for socioeconomic and demographic factors, providing a fairer measure of campus effectiveness.
Question 4: A teacher administers a five-question exit ticket at the end of each lesson and adjusts the next day's instruction based on results. This is an example of:
- Summative assessment driving curriculum mapping
- Formative assessment informing instructional decisions (Correct answer)
- Diagnostic assessment identifying learning disabilities
- Benchmark assessment predicting STAAR performance
Correct answer: Formative assessment informing instructional decisions
Brief checks for understanding used to adjust upcoming instruction are classic formative assessments embedded in the instructional cycle.
Question 5: Under the Texas accountability system, which metric specifically measures the academic growth of individual students from one year to the next?
- STAAR percent at Approaches Grade Level
- Student Growth component within the STAAR domain (Correct answer)
- Participation rate index
- Closing the Gaps domain score
Correct answer: Student Growth component within the STAAR domain
Texas's STAAR-based accountability system includes a Student Growth measure that tracks how individual student performance changes over consecutive years.
Question 6: A principal wants to measure whether the campus's new math intervention program caused student achievement gains. Which evaluation design provides the strongest causal evidence?
- A pre-post comparison using only intervention students
- A randomized controlled trial with a comparison group (Correct answer)
- Teacher surveys about perceived student improvement
- Analysis of district benchmark trends over three years
Correct answer: A randomized controlled trial with a comparison group
A randomized controlled trial (RCT) eliminates selection bias by randomly assigning students to treatment and control groups, providing the strongest causal evidence.
Question 7: A principal reviews campus assessment data and finds that reliability coefficients for the campus-created common assessments are below 0.70. What does this indicate?
- The assessments are too easy for the student population
- The assessments produce inconsistent results that limit their usefulness (Correct answer)
- The assessments are not aligned to TEKS objectives
- The assessments have too few answer choices per item
Correct answer: The assessments produce inconsistent results that limit their usefulness
A reliability coefficient below 0.70 signals that the assessment produces inconsistent scores, meaning results may reflect measurement error rather than true student knowledge.
A principal notices that students from one demographic group consistently score lower on a standardized test than other groups, despite similar classroom performance.
What should the principal investigate first?