โ† All Data Science Flashcard Decks

FREE Data Science Unsupervised Learning Techniques Questions and Answers Flashcards

6 cards from real Data Science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 FREE Data Science Unsupervised Learning Techniques Questions and Answers flashcards as text
  1. How do Gaussian Mixture Models differ from K-Means in their approach to clustering?

    Answer: They use probabilistic soft assignments based on multiple Gaussian distributions

    GMMs model each cluster as a Gaussian distribution and assign data points probabilistic memberships across all clusters rather than hard assignments.

  2. What is the main advantage of hierarchical agglomerative clustering over partitional methods?

    Answer: It produces a dendrogram showing cluster relationships at every level of granularity

    Agglomerative clustering builds a tree of merges (dendrogram) that lets analysts choose the number of clusters after seeing the full hierarchy.

  3. In Principal Component Analysis, what does the first principal component represent?

    Answer: The direction of maximum variance in the data

    The first principal component is the linear combination of original features that captures the greatest amount of variance in the dataset.

  4. What is a key limitation of using autoencoders for dimensionality reduction compared to PCA?

    Answer: Autoencoders can overfit and require careful regularization and tuning

    Autoencoders are neural networks that can memorize training data if not properly regularized, making them prone to overfitting on small datasets.

  5. Which evaluation metric for clustering does NOT require ground truth labels?

    Answer: Davies-Bouldin Index

    The Davies-Bouldin Index evaluates clustering quality using only within-cluster scatter and between-cluster distances, requiring no external labels.

  6. What is the primary goal of the Isolation Forest algorithm?

    Answer: Detecting anomalies by isolating observations through random partitioning

    Isolation Forest detects anomalies by randomly partitioning data and identifying points that require fewer splits to isolate, as outliers are easier to separate.