Unsupervised Learning: Clustering Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Unsupervised Learning: Clustering flashcards as text
What is the key difference between soft clustering and hard clustering?
Answer: Soft clustering assigns probabilities of membership to multiple clusters; hard clustering assigns each point to exactly one cluster
In soft (fuzzy) clustering like GMM, each point has a fractional membership probability across all clusters, unlike hard clustering where membership is binary.
When applying K-Means to text data represented as TF-IDF vectors, which distance metric is typically preferred?
Answer: Cosine similarity
Cosine similarity measures the angle between document vectors regardless of magnitude, making it more appropriate for high-dimensional sparse TF-IDF representations.
What does the 'elbow' in the K-Means elbow method represent?
Answer: The k value where marginal reduction in inertia becomes small
The elbow is where adding more clusters yields diminishing returns in inertia reduction, suggesting that additional clusters explain little additional variance.
OPTICS is an extension of DBSCAN that addresses which of its main limitations?
Answer: DBSCAN struggles with clusters of varying density
OPTICS produces a reachability plot that reveals clusters at multiple density levels, overcoming DBSCAN's single global epsilon that fails when clusters have different densities.
A customer segmentation model produces 6 clusters, but cluster 3 has a silhouette score of -0.2 while others average 0.6. What should the analyst do?
Answer: Try a different k value or merge cluster 3 with its nearest cluster
A negative silhouette score means points in cluster 3 are on average closer to another cluster than their own, indicating a poor cluster assignment that warrants reconfiguration.
In hierarchical clustering, what is the time complexity of the naive algorithm for n data points?
Answer: O(n³)
Naive agglomerative hierarchical clustering requires O(n²) space for the distance matrix and O(n³) time to repeatedly find and update minimum distances across n-1 merge steps.
Which statement best describes the concept of 'cohesion' in cluster evaluation?
Answer: How similar points within the same cluster are to each other
Cohesion measures intra-cluster compactness — high cohesion means points within a cluster are tightly grouped and similar to one another.