PAC Research Methodologies & Data Analysis 2 — Questions and Answers
Question 1: A policy analyst wants to determine whether a new job training program caused a reduction in unemployment. Which research design best establishes causality?
- Cross-sectional survey of current participants
- Randomized controlled trial with treatment and control groups (Correct answer)
- Case study of one successful participant
- Focus group with program administrators
Correct answer: Randomized controlled trial with treatment and control groups
Randomized controlled trials (RCTs) are the gold standard for establishing causality because random assignment eliminates selection bias.
Question 2: When conducting a difference-in-differences (DiD) analysis, what is the key assumption that must hold?
- Treatment and control groups must be identical in size
- Both groups would have followed parallel trends absent the intervention (Correct answer)
- The intervention must affect all units equally
- Pre-treatment data must span at least 10 years
Correct answer: Both groups would have followed parallel trends absent the intervention
The parallel trends assumption holds that in the absence of treatment, the average outcomes for treatment and control groups would have followed the same trend over time.
Question 3: A dataset on household income contains values ranging from $20,000 to $500,000 with several extreme outliers above $2 million. Which measure of central tendency best represents the typical household income?
- Mean
- Median (Correct answer)
- Mode
- Range
Correct answer: Median
The median is resistant to extreme outliers, making it the best measure of central tendency for skewed income distributions.
Question 4: In regression analysis, multicollinearity occurs when two or more independent variables are highly correlated. What is the primary problem this causes?
- It inflates the R-squared value artificially
- It makes coefficient estimates unstable and difficult to interpret (Correct answer)
- It causes the dependent variable to become categorical
- It eliminates statistical significance from all variables
Correct answer: It makes coefficient estimates unstable and difficult to interpret
Multicollinearity inflates the standard errors of correlated predictors, making individual coefficient estimates unreliable even if the overall model fits well.
Question 5: A policy researcher uses a Likert scale (1–5) to measure public satisfaction with government services. This type of data is best classified as:
- Nominal
- Ordinal (Correct answer)
- Interval
- Ratio
Correct answer: Ordinal
Likert scale data is ordinal because responses have a meaningful order, but the intervals between points are not necessarily equal.
Question 6: Which sampling method is most appropriate when a researcher needs to ensure that subgroups (e.g., urban/rural, income levels) are proportionally represented in the sample?
- Simple random sampling
- Convenience sampling
- Stratified random sampling (Correct answer)
- Snowball sampling
Correct answer: Stratified random sampling
Stratified random sampling divides the population into subgroups (strata) and randomly samples from each, ensuring proportional representation.
Question 7: A meta-analysis reviewing 20 studies on a policy's effectiveness finds high heterogeneity (I² = 85%). What does this indicate?
- The studies are nearly identical in their findings
- There is substantial variation in effect sizes across studies that cannot be ignored (Correct answer)
- The policy has no measurable effect
- The sample sizes across studies are inconsistent
Correct answer: There is substantial variation in effect sizes across studies that cannot be ignored
An I² of 85% indicates that 85% of the variability in effect sizes is due to true heterogeneity between studies, not sampling error, warranting a random-effects model.
A policy analyst wants to determine whether a new job training program caused a reduction in unemployment.
Which research design best establishes causality?