← All MS-DS Master of Data science Flashcard Decks

Machine Learning Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Machine Learning flashcards as text
  1. Which ensemble method trains each successive tree to correct the residual errors of the previous trees?

    Answer: Gradient Boosting

    Gradient Boosting fits each new learner to the residuals (negative gradients) of the ensemble so far, progressively reducing prediction error.

  2. In the bias-variance tradeoff, which condition describes a model that performs well on training data but poorly on unseen data?

    Answer: Low bias, high variance

    Low bias means the model fits training data well, but high variance means it captures noise and generalizes poorly to new data — the classic overfitting scenario.

  3. What is the primary purpose of the kernel trick in Support Vector Machines?

    Answer: To implicitly map data to a higher-dimensional space without computing the transformation explicitly

    The kernel trick computes dot products in a high-dimensional feature space implicitly, enabling SVMs to find nonlinear decision boundaries without the computational cost of explicit transformation.

  4. An ROC curve plots True Positive Rate against which metric?

    Answer: False Positive Rate

    The ROC curve plots TPR (sensitivity) on the y-axis versus FPR (1 − specificity) on the x-axis across all classification thresholds.

  5. Which regularization technique randomly drops neurons during training to prevent co-adaptation?

    Answer: Dropout

    Dropout randomly sets a fraction of neuron activations to zero during each training step, forcing the network to learn redundant representations.

  6. In k-means clustering, what does the algorithm minimize?

    Answer: Within-cluster sum of squared distances to centroids

    K-means minimizes the within-cluster sum of squared Euclidean distances (inertia) between each point and its assigned cluster centroid.

  7. Which activation function is most commonly used in the hidden layers of modern deep neural networks due to its ability to mitigate vanishing gradients?

    Answer: ReLU

    ReLU (Rectified Linear Unit) outputs zero for negative inputs and the identity for positive ones, providing non-saturation in the positive region and helping gradients flow during backpropagation.