← All AI Flashcard Decks

AI Engineer: Machine Learning Fundamentals and Algorithms Flashcards

7 cards from real AI practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 AI Engineer: Machine Learning Fundamentals and Algorithms flashcards as text
  1. What is the kernel trick in Support Vector Machines (SVM)?

    Answer: A technique that implicitly maps data to a higher-dimensional space to find a linear separator

    The kernel trick computes dot products in a higher-dimensional feature space without explicitly transforming the data, enabling SVMs to find linear decision boundaries for data that is nonlinearly separable in the original space.

  2. What is the primary difference between gradient boosting and bagging (e.g., Random Forest)?

    Answer: Gradient boosting builds models sequentially where each model corrects the errors of the previous one

    Gradient boosting builds an ensemble sequentially, where each new model focuses on correcting the residual errors of the combined previous models, reducing bias — unlike bagging which builds models independently in parallel to reduce variance.

  3. Why is feature standardization (zero mean, unit variance) important before applying algorithms like SVM or k-nearest neighbors?

    Answer: It prevents features with larger numerical ranges from dominating distance-based calculations

    Distance-based algorithms are sensitive to feature scale — a feature ranging 0–10,000 will dominate a feature ranging 0–1 in distance calculations, so standardization puts all features on equal footing.

  4. What does Principal Component Analysis (PCA) accomplish?

    Answer: It reduces dimensionality by projecting data onto orthogonal axes of maximum variance

    PCA finds orthogonal principal components (linear combinations of original features) ordered by the amount of variance they explain, allowing dimensionality reduction by keeping only the top components.

  5. In k-means clustering, how is the optimal number of clusters (k) typically determined?

    Answer: By using the elbow method, which plots inertia vs. k and identifies the point of diminishing returns

    The elbow method plots the within-cluster sum of squares (inertia) against different values of k — the 'elbow' point where adding more clusters yields diminishing reductions in inertia suggests the optimal k.

  6. What is the purpose of a validation set, distinct from both the training set and test set?

    Answer: To tune hyperparameters and select the best model without contaminating the final test evaluation

    The validation set is used during development to tune hyperparameters and compare models — using the test set for this purpose would cause leakage, making the test set an unreliable measure of real-world performance.

  7. Which of the following statements best describes a naive Bayes classifier?

    Answer: It applies Bayes' theorem assuming conditional independence between features given the class label

    Naive Bayes applies Bayes' theorem to compute the posterior probability of each class and classifies based on the highest probability, making the 'naive' assumption that features are conditionally independent given the class.