โ† All AML Flashcard Decks

Supervised & Unsupervised Learning Algorithms Flashcards

7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Supervised & Unsupervised Learning Algorithms flashcards as text
  1. Which loss function is minimized by logistic regression during training?

    Answer: Binary Cross-Entropy (Log Loss)

    Logistic regression minimizes binary cross-entropy (negative log-likelihood), which penalizes confident wrong predictions more heavily than uncertain ones.

  2. What is the key difference between agglomerative and divisive hierarchical clustering?

    Answer: Agglomerative starts with each point as its own cluster and merges; divisive starts with one cluster and splits

    Agglomerative clustering is bottom-up: each point starts as its own cluster, and pairs are merged iteratively; divisive is top-down, splitting clusters recursively.

  3. In a decision tree, information gain is computed as the reduction in which quantity after a split?

    Answer: Entropy

    Information gain measures the reduction in entropy (Shannon entropy) of the target variable achieved by splitting on a particular feature.

  4. What is the primary advantage of using Isolation Forest over LOF for anomaly detection?

    Answer: Isolation Forest scales efficiently to high-dimensional, large datasets

    Isolation Forest has linear time complexity and scales well to high dimensions and large datasets by isolating anomalies with random partitioning, unlike the quadratic-complexity LOF.

  5. Which scenario best illustrates the concept of 'label leakage' in a supervised classification task?

    Answer: Including a feature derived from the target variable in training features

    Label leakage occurs when a feature directly or indirectly encodes information about the target variable, causing artificially high training performance that won't generalize.

  6. In AdaBoost, what happens to the weights of incorrectly classified samples after each boosting round?

    Answer: Their weights are increased so the next classifier prioritizes them

    AdaBoost increases the sample weights of misclassified instances, forcing subsequent weak learners to focus more attention on the harder-to-classify examples.

  7. What does the linkage criterion 'Ward's method' minimize when merging clusters in hierarchical clustering?

    Answer: Increase in total within-cluster variance after the merge

    Ward's method merges the two clusters whose fusion causes the smallest increase in total within-cluster sum of squares, producing compact, roughly equal-sized clusters.