Research & Data Analysis Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Research & Data Analysis flashcards as text
A researcher wants to determine whether a new teaching method improves test scores compared to a traditional method. Which study design is most appropriate?
Answer: Randomized controlled experiment
A randomized controlled experiment allows causal inference by randomly assigning participants to treatment and control groups.
In hypothesis testing, a p-value of 0.03 with a significance level of 0.05 means:
Answer: Reject the null hypothesis
Since 0.03 < 0.05, the result is statistically significant and we reject the null hypothesis.
Which technique is used to assess the stability of a regression model by partitioning data into training and validation subsets multiple times?
Answer: K-fold cross-validation
K-fold cross-validation splits data into k subsets and iteratively trains/validates across all folds to estimate model performance.
A dataset has a mean of 50, median of 45, and mode of 40. This distribution is best described as:
Answer: Positively skewed
When mean > median > mode, the distribution has a longer right tail, indicating positive (right) skew.
A data scientist uses Lasso regression instead of OLS. The primary advantage of Lasso in high-dimensional data is:
Answer: It performs automatic feature selection by shrinking some coefficients to zero
Lasso's L1 penalty shrinks some coefficients to exactly zero, effectively selecting a sparse subset of predictors.
In a meta-analysis, publication bias most commonly results in:
Answer: Overestimation of effect sizes because significant results are more likely published
Studies with statistically significant results are more likely published, inflating the pooled effect size in meta-analyses.
Which measure of association is most appropriate when analyzing the relationship between two continuous variables that may not have a linear relationship?
Answer: Spearman rank correlation
Spearman rank correlation captures monotonic relationships and is robust to non-linearity and outliers unlike Pearson correlation.