Unsupervised Learning Techniques Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Unsupervised Learning Techniques flashcards as text
What is the 'manifold hypothesis' and why is it important for unsupervised learning?
Answer: It posits that high-dimensional data lies on a lower-dimensional manifold, motivating nonlinear dimensionality reduction
The manifold hypothesis states that real-world high-dimensional data concentrates near a low-dimensional curved surface, justifying methods like t-SNE, UMAP, and Isomap for dimensionality reduction.
Which hyperparameter in DBSCAN determines the minimum number of points required to form a dense region (core point)?
Answer: minPts
A point is a core point if at least minPts points (including itself) fall within its epsilon-neighborhood; minPts controls the density requirement for cluster formation.
When comparing GMMs to k-means, which advantage do GMMs offer?
Answer: GMMs can model clusters of varying shapes, sizes, and densities through covariance matrices
GMMs learn full covariance matrices per component, allowing elliptical clusters of different orientations and sizes, while k-means assumes spherical, equal-variance clusters.
A data scientist applies PCA to a dataset and finds the first two principal components explain 62% of variance. What is the most appropriate interpretation?
Answer: The 2D projection retains 62% of the information; the rest is lost in the low-dimensional representation
Explained variance ratio tells you how much information is retained in the projection; 62% means substantial structure is captured, but 38% is sacrificed in this 2D reduction.
What is the Calinski-Harabasz index, and what does a higher value indicate?
Answer: It is the ratio of between-cluster dispersion to within-cluster dispersion; higher values indicate better-defined clusters
The Calinski-Harabasz index (also called Variance Ratio Criterion) rewards tight, well-separated clusters: higher values mean clusters are dense internally and far from each other.
In the context of topic modeling, what does Latent Dirichlet Allocation (LDA) model as latent variables?
Answer: Per-document topic distributions and per-topic word distributions
LDA treats each document as a mixture of topics (θ) and each topic as a distribution over words (φ), inferring both as latent variables from observed word counts.
Which phenomenon explains why distance-based clustering methods can struggle with very high-dimensional data?
Answer: The curse of dimensionality causes distances between points to become increasingly uniform, reducing discriminative power
In high dimensions, the ratio of maximum to minimum pairwise distances converges to 1, making all points seem equidistant and undermining distance-based similarity measures used in clustering.