Unsupervised Learning: Clustering Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Unsupervised Learning: Clustering flashcards as text
What is the primary purpose of the 'gap statistic' method in clustering?
Answer: To compare within-cluster dispersion to a null reference distribution to find optimal k
The gap statistic compares WCSS of the actual clustering to that expected under a null uniform distribution, with the optimal k being where the gap is maximized.
In the context of clustering high-dimensional data, what is the 'curse of dimensionality' problem?
Answer: Distance metrics become less meaningful as all pairwise distances converge
As dimensionality increases, the ratio of maximum to minimum pairwise distances approaches 1, making distance-based clustering unreliable since all points appear equally far apart.
Which subspace clustering technique projects data into lower dimensions before clustering to mitigate high-dimensionality issues?
Answer: Applying PCA before K-Means
Applying PCA before K-Means reduces noise dimensions and retains variance-rich components, making distance computations more meaningful in the reduced space.
When using Fuzzy C-Means clustering, what does the fuzziness parameter 'm' control?
Answer: The degree of overlap between cluster memberships
Higher values of m (m > 1) increase fuzziness so memberships are spread more evenly across clusters; as m → 1, Fuzzy C-Means approaches hard K-Means partitioning.
A retail analyst clusters purchase transactions and finds that scaling features dramatically changes cluster assignments. Why does this happen with K-Means?
Answer: K-Means is sensitive to feature scale because it uses Euclidean distance
Euclidean distance gives more weight to features with larger numeric ranges, so an unscaled feature like 'transaction amount in dollars' dominates 'number of items bought'.
What is the main difference between divisive and agglomerative hierarchical clustering?
Answer: Divisive starts with one cluster and splits; agglomerative starts with n clusters and merges
Agglomerative (bottom-up) begins with each point as its own cluster and merges, while divisive (top-down) begins with all points in one cluster and recursively splits.
In which scenario would Gaussian Mixture Models be preferred over K-Means for customer segmentation?
Answer: When customers can belong to multiple segments with varying degrees
GMMs provide soft probabilistic memberships, making them ideal when customers can exhibit mixed behaviors (e.g., 70% value-seeker, 30% brand-loyal) rather than belonging to exactly one segment.