Research & Data Analysis Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Research & Data Analysis flashcards as text
A Likert-scale survey question (1–5) measuring satisfaction is best analyzed with which statistical approach?
Answer: Mann-Whitney U test as a non-parametric alternative
Ordinal Likert data often violates normality assumptions, making the Mann-Whitney U test a robust non-parametric choice for group comparisons.
In a multiple regression model, the variance inflation factor (VIF) is used to detect:
Answer: Multicollinearity among predictors
VIF quantifies how much a predictor's variance is inflated due to linear correlations with other predictors in the model.
Which sampling method ensures proportional representation of subgroups within the population?
Answer: Stratified random sampling
Stratified random sampling divides the population into strata and samples from each proportionally, guaranteeing subgroup representation.
A receiver operating characteristic (ROC) curve plots sensitivity against:
Answer: 1 – Specificity (false positive rate)
The ROC curve plots the true positive rate (sensitivity) on the y-axis against the false positive rate (1 – specificity) on the x-axis across all classification thresholds.
In Bayesian inference, the posterior distribution is proportional to:
Answer: The product of the prior and the likelihood
Bayes' theorem states posterior ∝ likelihood × prior; the marginal likelihood acts as a normalizing constant.
A researcher finds r = 0.92 between two variables but the relationship is explained entirely by a third variable. This is called:
Answer: Spurious correlation
A spurious correlation appears strong but disappears when a confounding variable is controlled, revealing no true direct relationship.
When building a predictive model, overfitting is best identified by:
Answer: High training accuracy and low test accuracy
Overfitting occurs when a model memorizes training data, yielding high in-sample performance but poor generalization to unseen test data.