Data Analysis & Statistical Methods Flashcards
7 cards from real CEA practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Data Analysis & Statistical Methods flashcards as text
In a multiple regression model, multicollinearity refers to:
Answer: High correlation among independent variables
Multicollinearity occurs when two or more independent variables are highly correlated, making it difficult to isolate their individual effects.
Which test is most appropriate for detecting autocorrelation in regression residuals?
Answer: Durbin-Watson test
The Durbin-Watson test specifically tests for first-order autocorrelation (serial correlation) in OLS regression residuals.
A Type II error in hypothesis testing occurs when:
Answer: A false null hypothesis is not rejected
A Type II error (false negative) occurs when we fail to reject a null hypothesis that is actually false.
In time-series analysis, a process is considered weakly stationary if:
Answer: Its mean and variance are time-invariant and its autocovariance depends only on lag
Weak (covariance) stationarity requires a constant mean, constant variance, and autocovariance that depends only on the lag, not on the specific time period.
The interquartile range (IQR) is preferred over standard deviation as a measure of spread when:
Answer: The dataset contains significant outliers
The IQR (Q3 โ Q1) is robust to outliers because it focuses on the middle 50% of the data, unlike standard deviation which is sensitive to extreme values.
In logistic regression, the output of the model before applying the sigmoid function is called:
Answer: Log-odds (logit)
Logistic regression models the log-odds (logit) of the probability of an event; the sigmoid function transforms this into a probability between 0 and 1.
Which criterion penalizes model complexity more heavily than the Akaike Information Criterion (AIC)?
Answer: Bayesian Information Criterion (BIC)
The BIC applies a larger penalty for additional parameters than AIC, especially as sample size grows, leading to selection of more parsimonious models.