← All DAC Flashcard Decks

Machine Learning & Predictive Analytics Flashcards

6 cards from real DAC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Machine Learning & Predictive Analytics flashcards as text
  1. Which machine learning algorithm is best suited for classifying emails as spam or not spam?

    Answer: Naïve Bayes

    Naïve Bayes is a probabilistic machine learning algorithm based on Bayes' theorem, commonly used for classification tasks, especially text classification. It is highly effective for classifying emails as spam or not spam because it can quickly learn from the presence or absence of specific words and features within an email. Its simplicity and efficiency make it a popular choice for this application.

  2. What is the purpose of cross-validation in machine learning?

    Answer: To prevent overfitting

    Cross-validation is a crucial technique in machine learning used to assess how well a model generalizes to new, unseen data. By training and testing the model on different subsets of the data, it helps detect and prevent overfitting. Overfitting occurs when a model learns the training data too well, including noise, and performs poorly on new data.

  3. Which metric is commonly used to evaluate the performance of a regression model?

    Answer: Mean Squared Error (MSE)

    Mean Squared Error (MSE) is a widely used metric to evaluate the performance of regression models. It calculates the average of the squared differences between the predicted values and the actual values. MSE quantifies the average magnitude of the errors, with lower values indicating a better fit of the model to the data.

  4. What is the main goal of feature scaling in machine learning?

    Answer: To make numerical features comparable

    The main goal of feature scaling in machine learning is to transform numerical features so they have a similar range or distribution. This prevents features with larger values from dominating the learning process, ensuring that all features contribute equally to the model's performance. It is particularly important for algorithms sensitive to feature magnitudes, such as K-Nearest Neighbors or Support Vector Machines.

  5. Which of the following is an unsupervised learning algorithm?

    Answer: K-Means Clustering

    K-Means Clustering is an unsupervised learning algorithm that groups data points into 'k' clusters based on their similarity. Unlike supervised learning, it does not require labeled data, making it ideal for discovering inherent structures or patterns within a dataset. It aims to partition data into clusters where each data point belongs to the cluster with the nearest mean.

  6. Which technique is commonly used to reduce dimensionality in machine learning?

    Answer: Principal Component Analysis (PCA)

    Principal Component Analysis (PCA) is a powerful statistical technique commonly used for dimensionality reduction in machine learning. It transforms high-dimensional data into a lower-dimensional representation while retaining most of the original variance. PCA helps simplify models, reduce computational cost, and mitigate the curse of dimensionality by identifying the most important features.