AI Engineer: Machine Learning Fundamentals and Algorithms Flashcards
7 cards from real AI practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 AI Engineer: Machine Learning Fundamentals and Algorithms flashcards as text
What is the curse of dimensionality and how does it affect machine learning models?
Answer: High-dimensional data requires exponentially more training examples to maintain statistical coverage of the feature space
As the number of dimensions increases, the volume of the feature space grows exponentially, making data increasingly sparse — meaning far more training examples are needed to learn reliable patterns.
What is transfer learning in the context of machine learning?
Answer: Reusing a model pre-trained on one task as a starting point for a different but related task
Transfer learning leverages representations learned by a model on a large dataset (e.g., ImageNet or large text corpora) and fine-tunes them for a new, related task where labeled data may be scarce.
What does the ROC-AUC score measure in a binary classification model?
Answer: The probability that the model ranks a randomly chosen positive example higher than a randomly chosen negative example
AUC (Area Under the ROC Curve) represents the probability that the model assigns a higher predicted probability to a random positive instance than to a random negative instance — an AUC of 1.0 is perfect, 0.5 is random.
Which of the following is an example of a generative machine learning model?
Answer: Variational Autoencoder (VAE)
A Variational Autoencoder is a generative model that learns a latent distribution of the training data and can sample from it to generate new data points, unlike discriminative models that only learn class boundaries.
What is early stopping as a regularization technique during model training?
Answer: Halting training when validation loss stops improving to prevent overfitting
Early stopping monitors validation loss during training and halts the process when it starts increasing, saving the model weights from the epoch with the best validation performance to prevent overfitting.
In the context of decision trees, what is information gain?
Answer: The reduction in entropy (uncertainty) in the target variable achieved by splitting on a given feature
Information gain measures how much a feature split reduces entropy (disorder) in the target variable — features with the highest information gain are chosen as split points to build a tree that separates classes most effectively.
What is the primary difference between a parametric and a non-parametric machine learning model?
Answer: Parametric models assume a fixed functional form with a set number of parameters, while non-parametric models grow in complexity with the training data
Parametric models (e.g., linear regression, logistic regression) summarize data with a fixed number of parameters regardless of dataset size, while non-parametric models (e.g., k-NN, kernel SVM) retain training data and grow in complexity with more data.