โ† All MS-DS Master of Data science Flashcard Decks

MS-DS Master of Data science Unsupervised Learning Techniques Questions and Answers Flashcards

6 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 MS-DS Master of Data science Unsupervised Learning Techniques Questions and Answers flashcards as text
  1. Which unsupervised learning technique reduces dimensionality by finding orthogonal axes that maximize variance in the data?

    Answer: Principal Component Analysis (PCA)

    PCA identifies principal components as orthogonal directions of maximum variance, enabling dimensionality reduction while preserving the most information.

  2. In DBSCAN, what is a 'core point'?

    Answer: A point with at least MinPts neighbors within radius epsilon

    DBSCAN defines a core point as one that has at least MinPts data points within its epsilon-radius neighborhood, forming the dense regions of clusters.

  3. What does the silhouette coefficient measure in clustering evaluation?

    Answer: How similar a point is to its own cluster compared to the nearest neighboring cluster

    The silhouette coefficient ranges from -1 to 1 and compares intra-cluster cohesion with nearest-cluster separation for each data point.

  4. Which technique is most appropriate for discovering hierarchical groupings without specifying the number of clusters in advance?

    Answer: Agglomerative hierarchical clustering

    Agglomerative hierarchical clustering builds a dendrogram by iteratively merging the closest clusters, allowing the analyst to choose the number of clusters after the fact.

  5. In a Gaussian Mixture Model (GMM), what algorithm is typically used to estimate the model parameters?

    Answer: Expectation-Maximization (EM)

    The EM algorithm iterates between computing soft cluster assignments (E-step) and updating the Gaussian parameters to maximize likelihood (M-step).

  6. What is the primary purpose of t-SNE in unsupervised learning?

    Answer: Visualizing high-dimensional data in two or three dimensions

    t-SNE is a nonlinear dimensionality reduction technique designed specifically for visualizing high-dimensional datasets in low-dimensional space while preserving local structure.