Machine Learning & Data Science Flashcards
7 cards from real CAIC practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Machine Learning & Data Science flashcards as text
Which algorithm finds the hyperplane that maximizes the margin between classes in a classification problem?
Answer: Support Vector Machine
SVMs find the decision boundary that maximizes the minimum distance (margin) to the nearest training points of each class (support vectors).
What does a learning curve that shows high training accuracy but low validation accuracy most likely indicate?
Answer: Overfitting
A large gap between training and validation performance indicates the model memorized training data but fails to generalize to unseen examples.
In the context of neural networks, what does 'vanishing gradient' refer to?
Answer: Gradients becoming extremely small in early layers, stalling learning
In deep networks, repeated multiplication of small gradients through many layers causes them to approach zero, making early layers learn very slowly.
Which data splitting strategy ensures that the proportion of each class is maintained in both train and test sets?
Answer: Stratified split
Stratified splitting preserves the class distribution of the full dataset in each resulting subset, critical for imbalanced classification problems.
What is the purpose of hyperparameter tuning using techniques like Grid Search or Random Search?
Answer: To find model configuration settings that optimize validation performance
Hyperparameter tuning systematically searches the space of model configurations (e.g., learning rate, depth) to find values that maximize held-out performance.
What distinguishes a generative model from a discriminative model in machine learning?
Answer: Generative models learn the joint distribution P(X,Y); discriminative models learn P(Y|X)
Generative models model how data is generated (joint distribution), while discriminative models learn decision boundaries directly from input to label.
Which technique addresses multicollinearity by adding the sum of squared coefficients as a penalty to the loss function?
Answer: Ridge (L2) regularization
Ridge regression penalizes large coefficients via their squared sum, shrinking correlated features' weights toward zero without eliminating them entirely.