Data Science Trivia Question and Answers Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Data Science Trivia Question and Answers flashcards as text
What does the bias-variance tradeoff describe in machine learning?
Answer: The balance between underfitting and overfitting a model
The bias-variance tradeoff describes how reducing model bias (underfitting) tends to increase variance (overfitting), and finding the right balance is key to good generalization.
Which dimensionality reduction technique finds the directions of maximum variance in data?
Answer: Principal Component Analysis (PCA)
PCA identifies the principal components, which are orthogonal axes that capture the maximum variance in the dataset, allowing effective dimensionality reduction.
What is a confusion matrix used for?
Answer: Evaluating classification model performance
A confusion matrix displays true positives, true negatives, false positives, and false negatives to evaluate how well a classification model performs across all classes.
In a random forest algorithm, what technique is used to create diverse decision trees?
Answer: Bagging with random feature selection
Random forests use bootstrap aggregating (bagging) combined with random feature selection at each split to build diverse, decorrelated decision trees that collectively reduce overfitting.
What does the term 'feature engineering' refer to in data science?
Answer: Creating new input variables from existing data to improve model performance
Feature engineering is the process of using domain knowledge to create, transform, or select input variables that help a machine learning model make better predictions.
Which evaluation metric is most appropriate for imbalanced classification datasets?
Answer: F1 Score
The F1 Score balances precision and recall, making it far more informative than accuracy when one class heavily outnumbers the other in an imbalanced dataset.