Advanced Machine Learning (AML) Certification — Questions and Answers
Question 1: What is the primary purpose of model explainability tools like SHAP in production ML systems?
- To speed up model inference at scale
- To compress model artifacts for faster deployment
- To explain individual predictions and ensure model transparency for stakeholders and auditors (Correct answer)
- To automate feature engineering pipelines
Correct answer: To explain individual predictions and ensure model transparency for stakeholders and auditors
SHAP attributes each feature's contribution to individual predictions, supporting transparency and regulatory compliance in production.
Question 2: Which dimensionality reduction technique projects data onto directions of maximum variance?
- LDA
- UMAP
- t-SNE
- PCA (Correct answer)
Correct answer: PCA
PCA (Principal Component Analysis) finds orthogonal axes of maximum variance to compress feature dimensions.
Question 3: What is target encoding for categorical variables?
- Encoding the target variable as integers
- Binarizing the target for classification tasks
- Encoding text features using TF-IDF
- Replacing each category value with the mean of the target variable for that category (Correct answer)
Correct answer: Replacing each category value with the mean of the target variable for that category
Target encoding substitutes each category with the average target value for that category, directly capturing predictive power.
Question 4: In graph neural networks (GNNs), what is the 'over-smoothing' problem that limits the number of layers?
- Node representations converging to indistinguishable vectors as more neighborhood aggregation layers are added (Correct answer)
- Gradient vanishing caused by deep message-passing stacks
- Inability to handle directed graphs with asymmetric edge weights
- Excessive memory usage from dense adjacency matrix storage
Correct answer: Node representations converging to indistinguishable vectors as more neighborhood aggregation layers are added
With many aggregation steps, all node features converge toward the same stationary distribution, losing local discriminative information.
Question 5: What is a feature store in an MLOps architecture?
- A tool for automated feature selection
- A replica of the training database used for inference
- A cloud storage bucket for raw datasets
- A centralized repository for storing, sharing, and serving precomputed features for ML models (Correct answer)
Correct answer: A centralized repository for storing, sharing, and serving precomputed features for ML models
Feature stores provide consistent, reusable features across training and serving pipelines, eliminating training-serving skew.
Question 6: What is a machine learning data pipeline?
- A visualization tool for feature distributions
- A method for hyperparameter tuning
- An automated sequence of data processing steps from ingestion to model-ready format (Correct answer)
- A storage system for training datasets
Correct answer: An automated sequence of data processing steps from ingestion to model-ready format
A data pipeline automates and chains data collection, cleaning, transformation, and feature engineering steps for reproducibility.
Question 7: What is the curse of dimensionality in machine learning?
- The problem of having too many training samples
- The computational cost of training deep neural networks
- The phenomenon where data becomes sparse as dimensions increase, degrading model performance (Correct answer)
- The difficulty of visualizing high-dimensional data
Correct answer: The phenomenon where data becomes sparse as dimensions increase, degrading model performance
As feature dimensions increase, data points become increasingly sparse, making distance-based algorithms less effective.
Question 8: What is semantic segmentation in computer vision?
- Ranking image regions by confidence score
- Detecting and localizing objects with bounding boxes
- Generating a text caption for an image
- Assigning a class label to every pixel in an image (Correct answer)
Correct answer: Assigning a class label to every pixel in an image
Semantic segmentation classifies each pixel of an image into a category, producing a pixel-wise class map of the entire scene.
Question 9: What is the 'deadly triad' in reinforcement learning, and why is it problematic?
- Speed, accuracy, and memory — hardware constraints prevent jointly optimizing all three
- Function approximation, bootstrapping, and off-policy training — their combination can cause divergence and instability (Correct answer)
- Reward, policy, and value function — optimizing all three simultaneously leads to conflicting gradients
- Exploration, exploitation, and generalization — they create oscillating reward signals
Correct answer: Function approximation, bootstrapping, and off-policy training — their combination can cause divergence and instability
The deadly triad describes how combining function approximation, bootstrapping (TD updates), and off-policy learning introduces compounding approximation errors that can cause Q-value estimates to diverge.
Question 10: Which feature engineering technique creates new features from products or powers of existing features?
- Imputation
- Feature scaling
- Polynomial feature expansion (Correct answer)
- Label encoding
Correct answer: Polynomial feature expansion
Polynomial feature expansion generates new features as products or powers of existing ones, enabling linear models to capture non-linear patterns.
Question 11: Which architecture is most commonly used for image-to-image tasks like semantic segmentation and medical image analysis?
- VGG-16
- Inception v3
- U-Net (encoder-decoder with skip connections) (Correct answer)
- ResNet
Correct answer: U-Net (encoder-decoder with skip connections)
U-Net uses a contracting encoder path and an expanding decoder path with skip connections that preserve spatial detail for precise pixel prediction.
Question 12: Which NLP task involves assigning a label (positive, negative, neutral) to a piece of text based on its emotional tone?
- Named Entity Recognition
- Machine Translation
- Coreference Resolution
- Sentiment Analysis (Correct answer)
Correct answer: Sentiment Analysis
Sentiment analysis classifies text according to the opinion or emotion expressed, commonly as positive, negative, or neutral.
Question 13: A machine learning engineer discovers their deployed fraud detection model has a significantly higher false positive rate for customers in low-income zip codes. What is the MOST ethically appropriate immediate action?
- Pause the model and notify stakeholders while investigating the root cause (Correct answer)
- Retrain the model with more data from those zip codes without informing leadership
- Accept the disparity as an inevitable artifact of the training data
- Document the finding and schedule a review for next quarter
Correct answer: Pause the model and notify stakeholders while investigating the root cause
Discovering a discriminatory disparity requires immediate escalation and pausing the harmful system while a proper investigation is conducted.
Question 14: What is the BLEU score used to measure in NLP?
- Quality of machine-generated text by comparing n-gram overlap with reference translations (Correct answer)
- Classification accuracy on text labels
- Number of out-of-vocabulary tokens in a corpus
- Semantic similarity between two sentence embeddings
Correct answer: Quality of machine-generated text by comparing n-gram overlap with reference translations
BLEU measures how many n-grams in a machine translation match those in one or more reference translations, normalized by length.
Question 15: In Kubernetes-based ML serving, what is the purpose of a Horizontal Pod Autoscaler (HPA) configured on a model serving deployment?
- It distributes requests across multiple model versions using weighted routing
- It automatically scales the number of inference pods up or down based on CPU/memory or custom metrics (Correct answer)
- It replicates model weights to all nodes for faster loading
- It enforces GPU resource quotas per namespace
Correct answer: It automatically scales the number of inference pods up or down based on CPU/memory or custom metrics
HPA monitors resource utilization or custom metrics (e.g., requests per second) and adjusts the pod replica count to handle inference load fluctuations.
Question 16: Which metric category is most critical when monitoring a deployed classification model in production?
- Prediction confidence distribution and accuracy on live data (Correct answer)
- Model file size on disk
- Training loss from the most recent training run
- Number of API calls per second
Correct answer: Prediction confidence distribution and accuracy on live data
Monitoring live prediction confidence and accuracy reveals model drift and performance degradation before it significantly impacts business outcomes.
Question 17: A Chief Data Officer (CDO) establishes a data lineage tracking system across all ML pipelines. The PRIMARY governance benefit of this is:
- Reducing model training compute costs
- Automatically detecting bias in model outputs
- Improving model prediction latency
- Enabling auditors to trace data origin, transformations, and usage for compliance (Correct answer)
Correct answer: Enabling auditors to trace data origin, transformations, and usage for compliance
Data lineage systems enable full traceability of how data flows from source to model, which is essential for regulatory audits, incident investigations, and governance accountability.
Question 18: What is the purpose of max pooling in a convolutional neural network?
- Normalize activations across the batch
- Reduce spatial dimensions while retaining dominant features (Correct answer)
- Introduce non-linearity into the network
- Increase the spatial resolution of feature maps
Correct answer: Reduce spatial dimensions while retaining dominant features
Max pooling downsamples feature maps by taking the maximum value in each pooling window, reducing size while preserving strong activations.
Question 19: What does the 'elbow method' evaluate when selecting the optimal number of clusters for K-Means?
- Within-cluster sum of squares (inertia) as a function of k (Correct answer)
- The gap statistic compared to a reference null distribution
- Between-cluster variance divided by total variance
- Silhouette coefficient as a function of cluster count
Correct answer: Within-cluster sum of squares (inertia) as a function of k
The elbow method plots within-cluster sum of squares (inertia) against k and selects the point where the rate of decrease sharply diminishes, forming an 'elbow.'
Question 20: In contrastive learning (e.g., SimCLR), what is the role of the projection head?
- To map representations to a space where the contrastive loss is applied, then discarded at downstream fine-tuning (Correct answer)
- To align positional encodings across augmented views
- To normalize representations before passing them to the backbone
- To produce final class predictions during inference
Correct answer: To map representations to a space where the contrastive loss is applied, then discarded at downstream fine-tuning
The projection head is a small MLP used only during pretraining to compute the contrastive loss; the backbone encoder is used for downstream tasks.
Question 21: In multi-objective hyperparameter optimization, what does the Pareto frontier represent?
- The region of hyperparameter space with highest variance
- The average performance weighted by objective importance
- The single best configuration across all objectives
- The set of configurations where no objective can be improved without degrading another (Correct answer)
Correct answer: The set of configurations where no objective can be improved without degrading another
The Pareto frontier contains all non-dominated solutions — configurations where improving one objective (e.g., accuracy) necessarily worsens another (e.g., inference latency).
Question 22: A compliance officer requests interpretability reports for a credit scoring model. Which approach best meets both technical and regulatory communication needs?
- Share only the model's overall accuracy metric
- Provide only the model's source code
- Claim the model is a black box and cannot be explained
- Generate SHAP-based feature importance reports translated into plain-language decision explanations aligned with regulatory requirements (Correct answer)
Correct answer: Generate SHAP-based feature importance reports translated into plain-language decision explanations aligned with regulatory requirements
SHAP-based explanations translated into plain language satisfy both technical interpretability needs and regulatory requirements for decision transparency.
Question 23: A model's training loss continues to decrease while validation loss increases after epoch 20. What does this report?
- Underfitting
- Overfitting (Correct answer)
- Data leakage
- Learning rate too low
Correct answer: Overfitting
Divergence between training and validation loss curves is the classic signature of overfitting — the model memorizes training data but fails to generalize.
Question 24: What is the primary function of a model monitoring dashboard in production?
- To display hyperparameter search results from training runs
- To track production metrics like prediction drift, latency, and data quality in real time (Correct answer)
- To visualize the model's internal architecture and weights
- To automate model retraining on a fixed schedule
Correct answer: To track production metrics like prediction drift, latency, and data quality in real time
Monitoring dashboards give operations teams real-time visibility into model health, enabling rapid detection of and response to performance degradation.
Question 25: What fundamentally distinguishes SARSA from Q-learning?
- SARSA uses neural networks while Q-learning uses lookup tables
- SARSA is on-policy and updates using the actual next action taken; Q-learning is off-policy and updates using the greedy max-Q action (Correct answer)
- SARSA targets continuous action spaces while Q-learning targets discrete ones
- SARSA requires an environment model while Q-learning does not
Correct answer: SARSA is on-policy and updates using the actual next action taken; Q-learning is off-policy and updates using the greedy max-Q action
SARSA (on-policy) updates Q-values using the action actually taken by the behavior policy, while Q-learning (off-policy) always bootstraps from the greedy maximum Q-value action.
Question 26: Which hyperparameter search strategy combines random search with intelligent sampling by building a density estimator over promising regions?
- Tree-structured Parzen Estimator (TPE) (Correct answer)
- Genetic algorithms
- Grid search with pruning
- Halving random search
Correct answer: Tree-structured Parzen Estimator (TPE)
TPE, used in Hyperopt, models the distribution of good and bad configurations separately using Parzen density estimators, directing search toward promising hyperparameter regions.
Question 27: Which tool is commonly used to containerize ML models for consistent deployment across environments?
- NumPy
- Scikit-learn
- Docker (Correct answer)
- Jupyter Notebook
Correct answer: Docker
Docker packages ML models with their dependencies into portable containers that run consistently across development and production environments.
Question 28: In AML practice, what is the purpose of a standard operating procedure (SOP)?
- To document step-by-step instructions for routine tasks to ensure consistency and quality (Correct answer)
- To restrict employee creativity
- To create unnecessary paperwork
- To satisfy management preferences only
Correct answer: To document step-by-step instructions for routine tasks to ensure consistency and quality
SOPs provide standardized, detailed instructions for routine operations, ensuring consistency, quality, efficiency, and safety regardless of which qualified individual performs the task.
Question 29: An ML team is evaluating two models: Model A achieves higher overall accuracy, while Model B has more equitable performance across demographic groups. According to fairness ethics, which factor should be the deciding consideration?
- Overall accuracy is a legally mandated criterion in the US
- Equity always takes precedence over accuracy regardless of context
- The deployment context and potential for disparate harm should guide the decision (Correct answer)
- Always choose the higher-accuracy model to maximize utility
Correct answer: The deployment context and potential for disparate harm should guide the decision
The appropriate fairness criterion depends on context — high-stakes decisions affecting vulnerable populations may require prioritizing equitable performance.
Question 30: When applying the EU AI Act's risk classification, an AI system used to evaluate job applicants' suitability would be classified as:
- High risk — requiring conformity assessment (Correct answer)
- Limited risk — requiring transparency obligations only
- Minimal risk — no specific requirements
- Unacceptable risk — banned outright
Correct answer: High risk — requiring conformity assessment
The EU AI Act Annex III classifies AI used in employment, worker management, and access to self-employment as high-risk, requiring conformity assessments and quality management systems.
Advanced Machine Learning (AML) Certification
The AML certification validates advanced competency across the full machine learning lifecycle, covering core algorithms, feature engineering, model deployment, NLP, computer vision, and professional skills including AI ethics, governance, and stakeholder communication.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds