← All MS-DS Master of Data science Flashcard Decks

Master of Data science Unsupervised Learning Techniques 1 Flashcards

6 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Master of Data science Unsupervised Learning Techniques 1 flashcards as text
  1. Which unsupervised learning algorithm assigns each data point to the nearest cluster centroid and iteratively updates centroids until convergence?

    Answer: K-Means

    K-Means iteratively assigns points to the nearest of K centroids (by Euclidean distance) and recomputes each centroid as the mean of its assigned points, repeating until assignments stop changing.

  2. In a Gaussian Mixture Model (GMM), the Expectation-Maximization (EM) algorithm alternates between which two steps?

    Answer: Computing posterior probabilities and updating distribution parameters

    The E-step computes the posterior probability (responsibility) that each Gaussian component generated each data point; the M-step updates the means, covariances, and mixture weights to maximize the expected log-likelihood.

  3. What is the primary purpose of t-SNE (t-Distributed Stochastic Neighbor Embedding)?

    Answer: Visualizing high-dimensional data in two or three dimensions while preserving local structure

    t-SNE converts pairwise similarities in high-dimensional space to probabilities and minimizes the KL divergence between those probabilities and ones in a low-dimensional embedding, making it ideal for 2-D or 3-D visualization of complex datasets.

  4. In hierarchical agglomerative clustering, which linkage criterion defines the distance between two clusters as the maximum pairwise distance between their members?

    Answer: Complete linkage

    Complete linkage uses the farthest (maximum) distance between any point in one cluster and any point in the other, producing compact, roughly equal-sized clusters and reducing the chaining effect seen with single linkage.

  5. An autoencoder trained without labels learns a compressed representation of input data in its bottleneck layer. What is this bottleneck representation called?

    Answer: Latent code (or latent space)

    The encoder compresses input into a lower-dimensional latent code; the decoder attempts to reconstruct the original input from it. The latent code captures the most salient features without any labeled supervision.

  6. Which metric is commonly used to determine the optimal number of clusters in K-Means by plotting within-cluster sum of squares against K and looking for a sharp bend?

    Answer: Elbow method

    The elbow method plots the within-cluster sum of squares (inertia) for increasing values of K; the point where the rate of decrease sharply slows—the 'elbow'—suggests the number of clusters beyond which adding more provides diminishing returns.