Unsupervised Learning Techniques Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Unsupervised Learning Techniques flashcards as text
Which linkage criterion in hierarchical clustering tends to produce the most compact, spherical clusters?
Answer: Ward's method
Ward's method minimizes the total within-cluster variance at each merge step, producing compact, roughly spherical clusters.
In the Expectation-Maximization (EM) algorithm for Gaussian Mixture Models, what does the E-step compute?
Answer: The posterior probability (responsibility) of each component for each data point
The E-step computes the responsibility r_{nk}, the posterior probability that component k generated data point n, given current parameters.
What is the primary role of the kernel function in kernel PCA?
Answer: To implicitly map data into a higher-dimensional feature space where linear PCA is applied
The kernel function computes inner products in a high-dimensional (possibly infinite) feature space without explicitly mapping points there, enabling nonlinear dimensionality reduction.
Which of the following is a key assumption made by the DBSCAN algorithm?
Answer: Clusters are defined by regions of high density separated by low-density regions
DBSCAN defines clusters as dense regions (cores + reachable points) separated by sparse regions, requiring only epsilon and minPts parameters.
In t-SNE, what does the perplexity hyperparameter control?
Answer: The balance between local and global structure by setting the effective number of neighbors
Perplexity roughly determines how many neighbors each point considers when constructing the high-dimensional probability distribution, balancing local vs. global structure.
What does the silhouette coefficient measure for a clustering result?
Answer: How similar a point is to its own cluster compared to neighboring clusters
The silhouette coefficient s(i) = (b(i) - a(i)) / max(a(i), b(i)), where a is intra-cluster distance and b is the nearest-cluster distance.
In Independent Component Analysis (ICA), what statistical property distinguishes the independent components from principal components?
Answer: ICA maximizes non-Gaussianity of components; PCA maximizes variance
ICA seeks statistically independent, non-Gaussian source signals, while PCA finds orthogonal directions of maximum variance (which are uncorrelated but not necessarily independent).