Unsupervised Learning Techniques Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Unsupervised Learning Techniques flashcards as text
What is the computational complexity of a single iteration of k-means on n data points with k clusters and d dimensions?
Answer: O(nkd)
Each of the n points must compute distances to each of the k cluster centroids in d dimensions, giving O(nkd) per iteration.
Which method is used by UMAP to preserve both local and global structure, distinguishing it from t-SNE?
Answer: UMAP constructs a fuzzy topological representation and optimizes its cross-entropy with a low-dimensional analog
UMAP is grounded in Riemannian geometry and algebraic topology, building a fuzzy simplicial complex and minimizing its cross-entropy with a low-dimensional representation, preserving more global structure than t-SNE.
What is the 'elbow method' used to determine in unsupervised learning?
Answer: The appropriate number of clusters k in k-means by plotting inertia vs. k
The elbow method plots within-cluster sum of squares (inertia) against k; the 'elbow' point where improvement diminishes suggests the optimal k.
A researcher applies hierarchical agglomerative clustering with single linkage to data with two elongated, touching clusters. What artifact is likely to occur?
Answer: Chaining: the two clusters will merge prematurely into one elongated cluster
Single linkage uses the minimum pairwise distance between clusters, making it susceptible to chaining where a bridge of close points merges two distinct elongated clusters early.
In a Variational Autoencoder (VAE), what is the purpose of the reparameterization trick?
Answer: To allow gradients to flow through the stochastic sampling step during backpropagation
The reparameterization trick expresses z = μ + σ·ε (ε ~ N(0,I)), making the random node deterministic w.r.t. ε and allowing gradient flow through μ and σ.
Which metric is specifically designed to evaluate clustering quality when ground-truth labels are available?
Answer: Adjusted Rand Index (ARI)
The Adjusted Rand Index measures the similarity between predicted cluster assignments and ground-truth labels, adjusted for chance agreement.
What distinguishes a self-organizing map (SOM) from k-means clustering?
Answer: SOM preserves topological relationships of data on a low-dimensional grid; k-means does not
SOMs arrange neurons on a 2D grid and update neighboring neurons during training, preserving the topological structure of the input space in the output map.