Supervised & Unsupervised Learning Algorithms Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Supervised & Unsupervised Learning Algorithms flashcards as text
What is the primary purpose of the 'kernel trick' in spectral clustering?
Answer: To implicitly compute similarities in high-dimensional feature spaces without explicit mapping
In spectral clustering, kernel functions allow computation of pairwise similarities in a high-dimensional space via the kernel trick, enabling non-linear cluster discovery.
In the context of Bayesian classifiers, what does the prior probability P(C) represent?
Answer: The probability of class C before observing any features
The prior P(C) encodes our belief about the relative frequency of class C in the population before any feature evidence is incorporated.
Which evaluation metric is most appropriate for an unsupervised clustering task when ground-truth labels are unavailable?
Answer: Silhouette Coefficient
The Silhouette Coefficient measures cohesion and separation using only the data's own structure, making it suitable when external ground-truth labels are not available.
What is the computational complexity of a single k-NN prediction for a dataset with N training samples and D features?
Answer: O(N * D)
A brute-force k-NN prediction requires computing the distance from the query point to all N training points, each requiring O(D) operations, giving O(N * D) total.
Which technique addresses the problem of imbalanced classes specifically during the training of a decision tree?
Answer: Using 'class_weight' parameter to penalize majority class errors more
Setting class weights inversely proportional to class frequencies causes the tree's impurity criterion to penalize misclassifying minority-class samples more, improving recall on rare classes.
What distinguishes a Restricted Boltzmann Machine (RBM) from a standard autoencoder as unsupervised representation learners?
Answer: RBMs are generative probabilistic models; autoencoders are deterministic encoder-decoder architectures
RBMs are energy-based generative models trained with contrastive divergence that learn a probability distribution over inputs, while autoencoders deterministically encode inputs to reconstruct them.
In stochastic gradient descent (SGD), why is shuffling the training data before each epoch important?
Answer: It prevents the model from memorizing the order of examples, reducing variance in gradient estimates
Shuffling prevents the optimizer from encountering systematically ordered examples that create correlated gradient updates, ensuring more uniform and unbiased parameter updates each epoch.