Unsupervised Learning Techniques Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Unsupervised Learning Techniques flashcards as text
What problem does the k-means++ initialization scheme address compared to random initialization?
Answer: It reduces the chance of poor convergence to local optima by spreading initial centroids
k-means++ selects each subsequent initial centroid with probability proportional to its squared distance from the nearest existing centroid, leading to better spread and faster convergence.
In spectral clustering, what is the graph Laplacian used for?
Answer: To embed data into a low-dimensional space where standard clustering is applied
Spectral clustering computes the eigenvectors of the graph Laplacian to embed the data into a space that reveals cluster structure, then applies k-means on this embedding.
Which of the following is a core limitation of PCA for dimensionality reduction in the context of unsupervised learning?
Answer: PCA only captures linear relationships and misses nonlinear structure in the data
PCA finds orthogonal directions of maximum variance using linear projections, so it cannot capture curved manifolds or other nonlinear structure in high-dimensional data.
A data scientist applies k-means to customer transaction data and observes that the algorithm assigns all points to one cluster in the first iteration. What is the most likely cause?
Answer: All initial centroids were placed in the same region of the feature space
If all centroids are initialized very close together, all points become nearest to one centroid, causing a degenerate solution; k-means++ initialization mitigates this.
In autoencoders used for anomaly detection, what serves as the anomaly score for a data point?
Answer: The reconstruction error between input and decoder output
An autoencoder trained on normal data learns to reconstruct normal patterns well; anomalies have high reconstruction error because the model cannot encode and decode them accurately.
Which statement best describes the difference between hard and soft clustering?
Answer: Hard clustering assigns each point to exactly one cluster; soft clustering assigns fractional memberships
In hard clustering (e.g., k-means), each point belongs to exactly one cluster; in soft clustering (e.g., GMM), each point has a probability of belonging to each cluster.
What is the primary difference between agglomerative and divisive hierarchical clustering?
Answer: Agglomerative starts with each point as its own cluster and merges; divisive starts with one cluster and splits
Agglomerative (bottom-up) begins with n singleton clusters and iteratively merges the most similar pair, while divisive (top-down) begins with all points in one cluster and recursively splits.