โ† All DSE Flashcard Decks

Knowledge Flashcards

7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Knowledge flashcards as text
  1. Which probability distribution is most appropriate for modeling the number of events occurring in a fixed time interval?

    Answer: Poisson distribution

    The Poisson distribution models the count of independent events occurring in a fixed interval of time or space.

  2. What does the term 'bias-variance tradeoff' describe in machine learning?

    Answer: The tension between underfitting (high bias) and overfitting (high variance)

    Bias-variance tradeoff refers to the tension between a model's error due to oversimplification (bias) and its sensitivity to training data fluctuations (variance).

  3. In SQL, what does a LEFT JOIN return?

    Answer: All rows from the left table and matching rows from the right

    A LEFT JOIN returns all rows from the left table and the matched rows from the right table; unmatched right-table rows appear as NULL.

  4. What is the purpose of cross-validation in model evaluation?

    Answer: To estimate model performance on unseen data using multiple train/test splits

    Cross-validation repeatedly splits data into training and validation sets to produce a more reliable estimate of generalization performance.

  5. Which measure of central tendency is most resistant to outliers?

    Answer: Median

    The median is resistant to outliers because it depends only on the middle value(s), not the magnitude of extreme values.

  6. What does PCA (Principal Component Analysis) primarily accomplish?

    Answer: Reduces dimensionality by projecting data onto directions of maximum variance

    PCA transforms features into a smaller set of uncorrelated principal components that capture the most variance in the data.

  7. In the context of hypothesis testing, what does a p-value represent?

    Answer: The probability of observing results at least as extreme as those seen, assuming the null hypothesis is true

    The p-value is the probability of obtaining a test statistic as extreme or more extreme than observed, under the assumption that the null hypothesis is true.