Supervised Learning: Classification Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Supervised Learning: Classification flashcards as text
What is 'concept drift' in the context of a deployed classification model?
Answer: The statistical properties of the target variable change over time, degrading model performance
Concept drift occurs when the relationship between input features and the target class shifts over time in production, causing model accuracy to degrade.
In a confusion matrix for binary classification, what does a False Positive (FP) represent?
Answer: A negative instance incorrectly classified as positive
A False Positive occurs when the model predicts the positive class but the true label is negative.
What is the purpose of calibrating a classifier's predicted probabilities?
Answer: To ensure the predicted probabilities accurately reflect true likelihoods
Calibration aligns predicted probabilities with actual observed frequencies, so a model predicting 0.8 is correct about 80% of the time.
Which ensemble method trains base classifiers sequentially, with each model focusing more on previously misclassified examples?
Answer: Boosting
Boosting trains classifiers sequentially, reweighting training samples so subsequent models focus on the errors of prior ones.
What is the 'curse of dimensionality' and how does it affect k-NN classification?
Answer: As dimensions increase, distances between points become similar, making nearest neighbors less meaningful
In high-dimensional spaces, Euclidean distances concentrate, making it hard to distinguish 'near' from 'far' neighbors, which undermines k-NN's core assumption.
Which of the following activation functions is used in the output layer of a neural network for binary classification?
Answer: Sigmoid
The sigmoid function maps the output to a value between 0 and 1, which can be interpreted as the probability of the positive class.
A model trained on historical loan data performs well on training data but poorly on new applicants. Which problem does this most likely indicate?
Answer: Overfitting due to high variance
Strong training performance but poor generalization to new data is the classic symptom of overfitting, where the model has learned noise or patterns specific to the training set.