โ† All Data Science Flashcard Decks

Data Science Supervised Learning Models Questions and Answers Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Data Science Supervised Learning Models Questions and Answers flashcards as text
  1. Which supervised learning method models the probability of a binary outcome using a logistic (sigmoid) function?

    Answer: Logistic Regression

    Logistic regression applies the sigmoid function to a linear combination of features to output probabilities for binary classification.

  2. What is the bias-variance tradeoff in supervised learning?

    Answer: The balance between a model's ability to fit training data closely and its ability to generalize to new data

    The bias-variance tradeoff describes how reducing bias (underfitting) often increases variance (overfitting) and vice versa.

  3. In a k-Nearest Neighbors classifier, what effect does increasing the value of k typically have on the decision boundary?

    Answer: The decision boundary becomes smoother and less sensitive to noise

    A larger k averages over more neighbors, which smooths the decision boundary and reduces sensitivity to individual noisy data points.

  4. Which technique splits the dataset into multiple folds to train and validate a supervised model, ensuring every observation is used for both?

    Answer: K-Fold Cross-Validation

    K-fold cross-validation divides data into k equal folds, rotating which fold serves as the validation set so every sample is tested exactly once.

  5. What assumption does Naive Bayes make about the relationship between features given the class label?

    Answer: All features are conditionally independent given the class

    Naive Bayes assumes that each feature contributes independently to the probability of a class, which simplifies computation but rarely holds exactly.

  6. What is the purpose of the cost function (loss function) in training a supervised learning model?

    Answer: To quantify the difference between predicted and actual values so the model can be optimized

    The cost function measures prediction error, and the training algorithm minimizes it to find optimal model parameters.