Unsupervised Learning: Clustering Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Unsupervised Learning: Clustering flashcards as text
Which internal cluster validation metric measures the ratio of intra-cluster distances to inter-cluster distances?
Answer: Davies-Bouldin Index
The Davies-Bouldin Index measures the average similarity ratio of each cluster to its most similar cluster, where lower values indicate better clustering.
In DBSCAN, a point that is reachable from a core point but is not itself a core point is called a:
Answer: Border point
Border points lie within the epsilon neighborhood of a core point but do not have enough neighbors to be core points themselves.
What is the primary advantage of using Mini-Batch K-Means over standard K-Means?
Answer: Faster convergence on large datasets
Mini-Batch K-Means uses small random subsets of data in each iteration, dramatically reducing computation time while maintaining comparable quality.
Which linkage criterion in hierarchical clustering tends to produce compact, spherical clusters?
Answer: Ward linkage
Ward linkage minimizes the total within-cluster variance at each merge step, consistently producing compact and similarly-sized clusters.
A data scientist runs K-Means with k=5 and observes that one cluster has only 2 points out of 10,000. What is the most likely issue?
Answer: Poor initialization of centroids
Poor centroid initialization can trap a centroid in a sparse region, resulting in a near-empty cluster while others are overcrowded.
In spectral clustering, what is the role of the Laplacian matrix?
Answer: Encodes graph connectivity for eigendecomposition
The graph Laplacian captures the connectivity structure of the data, and its eigenvectors reveal cluster membership in the embedded space.
Which of the following clustering algorithms is best suited for discovering clusters of arbitrary shape in spatial data?
Answer: DBSCAN
DBSCAN defines clusters by density rather than distance to centroids, enabling it to find clusters of any shape including rings and crescents.