โ† All MS-DS Master of Data science Flashcard Decks

Unsupervised Machine Learning Models Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Unsupervised Machine Learning Models flashcards as text
  1. In Non-negative Matrix Factorization (NMF), what constraint separates it from standard matrix factorization methods?

    Answer: All elements in the factor matrices must be non-negative, enabling parts-based representations

    NMF constrains all factor values to be โ‰ฅ 0, which leads to additive, parts-based decompositions (e.g., facial parts instead of eigenfaces).

  2. Which distance metric is most appropriate when comparing high-dimensional sparse vectors such as TF-IDF document representations?

    Answer: Cosine similarity

    Cosine similarity measures the angle between vectors and is invariant to document length, making it ideal for sparse high-dimensional text data.

  3. What is the primary purpose of the encoder in a Variational Autoencoder (VAE) compared to a standard autoencoder?

    Answer: The VAE encoder outputs parameters of a probability distribution rather than a fixed latent vector

    The VAE encoder predicts a mean and variance, enabling sampling from a learned latent distribution and smooth interpolation in latent space.

  4. In spectral clustering, what is the role of the graph Laplacian?

    Answer: It encodes pairwise connectivity, and its eigenvectors reveal the cluster structure of the data

    The eigenvectors of the graph Laplacian capture the connectivity structure of the similarity graph, allowing k-means to be applied in the spectral embedding space.

  5. Which criterion is commonly used to select the optimal number of topics in LDA?

    Answer: Perplexity on a held-out test corpus

    LDA perplexity measures how well the trained model predicts unseen documents; lower perplexity on held-out data indicates a better model fit.

  6. What is 'concept drift' and how does it challenge unsupervised streaming clustering?

    Answer: The underlying data distribution changes over time, invalidating previously learned cluster structures

    Concept drift occurs when the statistical properties of the input stream evolve, requiring streaming clustering algorithms to adapt their cluster assignments dynamically.

  7. What does the 'curse of dimensionality' imply specifically for distance-based clustering algorithms?

    Answer: In high dimensions, distances between points become nearly equal, making separation of clusters difficult

    As dimensionality increases, the ratio of maximum to minimum pairwise distances approaches 1, making it hard to identify meaningful nearest neighbors for clustering.

Unsupervised Machine Learning Models Flashcards โ€” MS-DS Master of Data science Study Cards with Answers