Machine Learning Fundamentals Flashcards
7 cards from real Data and Analytics practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Machine Learning Fundamentals flashcards as text
Which metric is most appropriate when evaluating a classification model on a highly imbalanced dataset?
Answer: F1-score
F1-score balances precision and recall, making it a better metric than accuracy for imbalanced datasets where the majority class can dominate the accuracy score.
What does an ROC curve illustrate in model evaluation?
Answer: The tradeoff between true positive rate and false positive rate at various thresholds
An ROC curve plots the true positive rate against the false positive rate across different classification thresholds, showing a model's discriminative ability.
In k-fold cross-validation, what happens to the data?
Answer: The data is divided into k equal parts; each fold serves as the test set once while the rest train the model
K-fold cross-validation partitions data into k folds, cycling through each fold as the test set while training on the remaining k-1 folds, then averaging the results.
What is a confusion matrix used to measure?
Answer: The counts of true positives, true negatives, false positives, and false negatives
A confusion matrix displays the counts of correct and incorrect predictions broken down by class, enabling calculation of precision, recall, and other classification metrics.
Which of the following best describes hyperparameter tuning?
Answer: Searching for the best configuration settings that control the learning process
Hyperparameter tuning involves searching for optimal settings (like learning rate, tree depth, or regularization strength) that are set before training and control how the model learns.
What is the primary advantage of using ensemble methods like bagging?
Answer: They combine multiple models to reduce variance and improve predictive performance
Bagging (Bootstrap Aggregating) trains multiple models on different bootstrap samples and averages their predictions, reducing variance and improving overall model stability.
Which approach is used to handle missing values in a dataset by replacing them with the mean, median, or mode?
Answer: Imputation
Imputation fills in missing values with a statistical estimate (mean, median, or mode) or a model-based prediction to preserve data completeness for training.