Data Science Unsupervised Learning Techniques Questions and Answers Flashcards
6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Data Science Unsupervised Learning Techniques Questions and Answers flashcards as text
Which unsupervised learning technique reduces the dimensionality of data by finding principal components that maximize variance?
Answer: Principal Component Analysis (PCA)
PCA identifies orthogonal axes (principal components) that capture the maximum variance in the dataset, effectively reducing dimensionality.
In DBSCAN clustering, what is a 'core point'?
Answer: A point with at least minPts neighbors within epsilon distance
A core point in DBSCAN is defined as a point that has at least minPts data points within its epsilon-radius neighborhood.
What does the silhouette coefficient measure in cluster analysis?
Answer: How similar a point is to its own cluster compared to neighboring clusters
The silhouette coefficient ranges from -1 to 1 and measures how well each data point fits within its assigned cluster versus the nearest neighboring cluster.
Which unsupervised technique is most appropriate for detecting unusual transactions in credit card data without labeled fraud examples?
Answer: Anomaly detection using Isolation Forest
Isolation Forest is an unsupervised anomaly detection method that isolates outliers by randomly partitioning data, making it ideal for fraud detection without labels.
What is the primary purpose of t-SNE in unsupervised learning?
Answer: Visualizing high-dimensional data in 2D or 3D space
t-SNE (t-distributed Stochastic Neighbor Embedding) is a nonlinear dimensionality reduction technique primarily used for visualizing complex high-dimensional datasets.
In hierarchical agglomerative clustering, what does 'complete linkage' use to measure distance between two clusters?
Answer: The maximum distance between any pair of points from the two clusters
Complete linkage defines inter-cluster distance as the maximum distance between any single pair of points from each cluster, tending to produce compact clusters.