Supervised & Unsupervised Learning Algorithms Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Supervised & Unsupervised Learning Algorithms flashcards as text
Which regularization technique used in neural networks randomly sets a fraction of neuron activations to zero during training?
Answer: Dropout
Dropout randomly zeroes a proportion of neuron outputs each forward pass, preventing co-adaptation of neurons and acting as an ensemble of many subnetworks.
In Principal Component Analysis (PCA), what do the eigenvalues of the covariance matrix represent?
Answer: The proportion of variance explained by each principal component
Each eigenvalue quantifies the amount of variance in the data that is captured by its corresponding eigenvector (principal component direction).
What distinguishes hard margin SVM from soft margin SVM?
Answer: Hard margin SVM allows no misclassifications; soft margin SVM allows controlled misclassifications via slack variables
Hard margin SVM requires all training points to be correctly classified and outside the margin, while soft margin SVM introduces slack variables to permit some violations for non-separable data.
Which statement about t-SNE (t-Distributed Stochastic Neighbor Embedding) is TRUE?
Answer: t-SNE uses Student's t-distribution in low-dimensional space to alleviate the crowding problem
t-SNE uses a heavy-tailed t-distribution in the low-dimensional map, which allows moderately dissimilar points to be placed further apart and alleviates the crowding problem.
When would you prefer Naive Bayes over Logistic Regression for text classification?
Answer: When training data is scarce and the conditional independence assumption is approximately met
Naive Bayes performs well with small datasets because it estimates feature probabilities independently, requiring far fewer parameters than logistic regression to learn.
In XGBoost, what role does the 'gamma' (min_split_loss) hyperparameter play?
Answer: It specifies the minimum loss reduction required to make a further partition on a leaf node
Gamma defines the minimum gain needed to justify splitting a leaf node; higher values make the algorithm more conservative and reduce overfitting by pruning unnecessary splits.
Which linkage method in hierarchical clustering is most sensitive to outliers?
Answer: Single linkage
Single linkage defines cluster distance as the minimum distance between any two points across clusters, making it highly sensitive to outliers that can form long 'chaining' clusters.