← All CAIC Flashcard Decks

Machine Learning & Data Science Flashcards

7 cards from real CAIC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Machine Learning & Data Science flashcards as text
  1. What is the purpose of the 'train-test split' in machine learning model development?

    Answer: To evaluate model generalization on data it was not trained on

    Holding out a test set that the model never sees during training provides an unbiased estimate of how well the model will perform on new data.

  2. In recommendation systems, what does 'collaborative filtering' rely on to make predictions?

    Answer: User behavior patterns and similarities across users or items

    Collaborative filtering identifies users with similar preferences or items with similar interaction patterns to predict what a user might like.

  3. Which activation function is most commonly used in hidden layers of modern deep neural networks due to its resistance to vanishing gradients?

    Answer: ReLU

    ReLU (Rectified Linear Unit) outputs zero for negative inputs and the input itself for positive inputs, avoiding saturation and enabling faster convergence.

  4. What does 'data leakage' mean in a machine learning context?

    Answer: Information from the test set influencing the training process

    Data leakage occurs when information that would not be available at prediction time is inadvertently included in training, inflating evaluation metrics.

  5. Which unsupervised learning method represents each data point by the distance to its nearest cluster center and detects anomalies as points far from all centers?

    Answer: K-Means anomaly detection

    In k-means-based anomaly detection, points with high distance to their assigned centroid are flagged as outliers since they don't fit any cluster well.

  6. What is the difference between a parameter and a hyperparameter in machine learning?

    Answer: Parameters are learned from training data; hyperparameters are set before training

    Parameters (e.g., weights, biases) are optimized during training, while hyperparameters (e.g., learning rate, tree depth) are configuration choices made prior to training.

  7. Which technique improves model performance on minority classes by synthetically creating new training examples?

    Answer: SMOTE (Synthetic Minority Over-sampling Technique)

    SMOTE generates synthetic minority-class samples by interpolating between existing minority instances, helping classifiers learn better decision boundaries.