โ† All MS-DS Master of Data science Flashcard Decks

Unsupervised Machine Learning Models Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Unsupervised Machine Learning Models flashcards as text
  1. When applying k-means to categorical data, which variant is the appropriate alternative?

    Answer: k-modes, which uses mode instead of mean and Hamming-based dissimilarity

    k-modes replaces the mean with the mode for categorical variables and uses a dissimilarity measure like simple matching, making it suitable for categorical data.

  2. In matrix factorization for collaborative filtering, what do the latent factors typically represent?

    Answer: Latent features capturing abstract user preferences and item attributes that explain observed ratings

    Latent factors are learned abstract representations (e.g., genre preference, quality sensitivity) that explain the pattern of observed ratings without explicit feature engineering.

  3. What is the 'elbow method' limitation when used to select k for k-means clustering?

    Answer: The elbow is often ambiguous or absent, making objective k selection difficult in practice

    Real-world inertia curves often decrease smoothly without a clear inflection point, making the elbow subjective and unreliable for automated k selection.

  4. Isolation Forest detects anomalies by:

    Answer: Measuring how few random splits are needed to isolate a point in a random tree ensemble

    Anomalies are isolated quickly (few splits) because they are sparse and distinct, while normal points require many splits to isolate in randomly partitioned trees.

  5. In the context of autoencoders used for anomaly detection, a high reconstruction error for a test sample indicates:

    Answer: The sample differs significantly from the distribution the autoencoder was trained on, suggesting an anomaly

    Autoencoders trained on normal data learn to reconstruct normal patterns well; anomalous inputs lie off-manifold and result in high reconstruction error.

  6. Which statement correctly describes the difference between density-based and centroid-based clustering?

    Answer: Density-based methods (e.g., DBSCAN) can find arbitrarily shaped clusters and handle noise, while centroid-based (e.g., k-means) assumes convex, roughly equal-sized clusters

    DBSCAN and similar methods group dense regions regardless of shape and label sparse points as noise, whereas k-means partitions all points into k convex Voronoi regions.

  7. What is the primary purpose of whitening as a preprocessing step before applying ICA?

    Answer: To decorrelate the data and normalize variances so ICA only needs to find rotations maximizing non-Gaussianity

    Whitening (sphering) decorrelates observations and scales them to unit variance, reducing the ICA problem to finding an orthogonal rotation that maximizes statistical independence.