Web Traffic Research & Evidence-Based Practice 2 — Questions and Answers
Question 1: Which statistical method is best suited for identifying which traffic channels contribute most to conversions in a multi-touch attribution study?
- Logistic regression
- Shapley value analysis (Correct answer)
- Chi-square test
- Moving average
Correct answer: Shapley value analysis
Shapley value analysis distributes conversion credit fairly across all touchpoints by evaluating each channel's marginal contribution.
Question 2: A researcher wants to determine whether a site redesign caused a drop in organic traffic or if the change was due to a Google algorithm update. Which approach best isolates the cause?
- Compare traffic before and after redesign only
- Run a regression discontinuity design using the redesign launch date as the cutoff (Correct answer)
- Survey users about the redesign
- Check bounce rate trends
Correct answer: Run a regression discontinuity design using the redesign launch date as the cutoff
Regression discontinuity exploits the sharp launch date as a natural experiment boundary to separate the redesign effect from concurrent algorithm changes.
Question 3: When analyzing web traffic data, what does a high variance inflation factor (VIF) indicate about your regression model?
- Strong predictive power of the model
- Multicollinearity among predictor variables (Correct answer)
- Heteroscedasticity in residuals
- Overfitting due to too many observations
Correct answer: Multicollinearity among predictor variables
A high VIF (typically >10) signals that predictor variables are highly correlated, making it difficult to isolate individual effects on traffic.
Question 4: Which research design is most appropriate for studying whether email newsletters cause repeat site visits without random assignment?
- Randomized controlled trial
- Difference-in-differences using subscribers vs. non-subscribers (Correct answer)
- Cross-sectional survey
- Case study
Correct answer: Difference-in-differences using subscribers vs. non-subscribers
Difference-in-differences compares the change in visits for subscribers vs. non-subscribers before and after newsletter campaigns, controlling for pre-existing differences.
Question 5: A content team publishes new blog posts every Monday. How should a researcher account for this systematic pattern when modeling weekly traffic fluctuations?
- Remove all Monday data points
- Add a day-of-week dummy variable to the regression (Correct answer)
- Use only weekend data for baseline
- Apply a Box-Cox transformation to the traffic variable
Correct answer: Add a day-of-week dummy variable to the regression
Day-of-week dummy variables capture systematic weekday effects, preventing publication day spikes from biasing the model coefficients.
Question 6: What is the primary advantage of using synthetic control methods over standard difference-in-differences in web traffic research?
- It requires a larger sample size
- It constructs a data-driven counterfactual when no single control unit exists (Correct answer)
- It eliminates the need for pre-treatment data
- It only works for binary outcomes
Correct answer: It constructs a data-driven counterfactual when no single control unit exists
Synthetic control builds a weighted combination of control units to approximate what the treated site's traffic would have been without the intervention.
Question 7: Which metric best captures the practical significance of a statistically significant increase in organic traffic?
- p-value
- Cohen's d effect size (Correct answer)
- Standard error
- Degrees of freedom
Correct answer: Cohen's d effect size
Cohen's d quantifies the magnitude of the traffic change in standard deviation units, distinguishing trivial from meaningful improvements regardless of sample size.
Which statistical method is best suited for identifying which traffic channels contribute most to conversions in a multi-touch attribution study?