โ† All MS-DS Master of Data science Flashcard Decks

Statistical and Probabilistic Analysis Flashcards

6 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Statistical and Probabilistic Analysis flashcards as text
  1. A data science team is conducting an A/B test on a new website feature. After collecting data, they calculate a p-value of 0.03. If their significance level (alpha) is set at 0.05, which of the following is the most accurate interpretation of the result?

    Answer: The result is statistically significant, meaning there is sufficient evidence to reject the null hypothesis.

    The p-value represents the probability of observing the collected data, or more extreme data, assuming the null hypothesis is true. [6, 21] Since the p-value (0.03) is less than the significance level (alpha = 0.05), the result is considered statistically significant. This provides enough evidence to reject the null hypothesis, which typically states there is no effect or difference. [12]

  2. A data scientist is working with a dataset representing the daily number of users visiting a specific webpage. The distribution of this data is highly skewed to the right. To perform hypothesis testing on the average number of daily users, which statistical principle allows the use of a Z-test or t-test, provided the sample size is large enough?

    Answer: The Central Limit Theorem

    The Central Limit Theorem (CLT) states that the distribution of sample means will be approximately normal, regardless of the population's original distribution, as long as the sample size is sufficiently large (often cited as n > 30). [2, 3, 23] This allows for the use of parametric tests like Z-tests or t-tests, which assume normality of the sampling distribution. [25]

  3. Which of the following scenarios is best modeled by a Poisson distribution?

    Answer: The number of customer service calls received by a call center in one hour.

    A Poisson distribution is used to model the number of events occurring within a fixed interval of time or space, given a known constant mean rate and independence of events. The number of calls per hour fits this description. [1, 15] A Binomial distribution is better for the coin flips and defective items (fixed number of trials). [14, 21] An Exponential distribution would model the time *between* events, not the count of events. [4]

  4. In the context of statistical analysis, which statement best describes the fundamental difference between the Frequentist and Bayesian approaches?

    Answer: Frequentists use probability to describe long-run frequencies of events, while Bayesians use it to quantify the degree of belief in a statement. [22]

    The core philosophical difference lies in the interpretation of probability. Frequentists view probability as the long-run frequency of an outcome in repeated experiments, treating population parameters as fixed, unknown constants. [9, 22] Bayesians, in contrast, use probability to represent a degree of belief or confidence about a parameter, which can be updated as more data becomes available. They treat parameters as random variables. [13, 16]

  5. A machine learning model is built to classify emails as 'spam' or 'not spam'. The model has a 99% accuracy rate. However, only 0.5% of all emails are actually spam. Why might accuracy be a misleading metric in this scenario?

    Answer: A high accuracy can be achieved by a naive model that classifies every email as 'not spam'.

    This is a classic example of a class imbalance problem. If a model simply classifies all emails as 'not spam', it would be correct 99.5% of the time, achieving high accuracy but completely failing at its goal of identifying spam. In such cases, other metrics like Precision, Recall, or the F1-score provide a much better assessment of the model's performance on the minority class.

  6. A data scientist calculates a 95% confidence interval for the average user engagement time on a new app feature to be [15.2 minutes, 18.6 minutes]. What is the correct interpretation of this interval?

    Answer: If we were to repeat this sampling process many times, 95% of the calculated confidence intervals would contain the true population mean.

    A confidence interval's interpretation is based on the long-run frequency of the procedure. The 95% confidence level means that if the same sampling method were used to generate many different samples, we would expect 95% of the resulting confidence intervals to capture the true, unknown population mean. [2] It does not assign a probability to the true parameter being within a specific, calculated interval. [6]

Statistical and Probabilistic Analysis Flashcards โ€” MS-DS Master of Data science Study Cards with Answers