Advanced Machine Learning (AML) Certification — Questions and Answers
Question 1: What is the purpose of max pooling in a convolutional neural network?
- Reduce spatial dimensions while retaining dominant features (Correct answer)
- Introduce non-linearity into the network
- Normalize activations across the batch
- Increase the spatial resolution of feature maps
Correct answer: Reduce spatial dimensions while retaining dominant features
Max pooling downsamples feature maps by taking the maximum value in each pooling window, reducing size while preserving strong activations.
Question 2: In AML practice, what is a needs assessment?
- An employee benefits survey
- A real estate appraisal
- A list of office supply requirements
- A systematic process to identify gaps between current conditions and desired outcomes (Correct answer)
Correct answer: A systematic process to identify gaps between current conditions and desired outcomes
A needs assessment systematically identifies and evaluates the gaps between current performance or conditions and desired outcomes, providing data-driven justification for programs, training, or interventions.
Question 3: Which outlier detection method uses the interquartile range (IQR) to flag extreme values?
- Min-max scaling
- Z-score normalization
- Winsorization
- Tukey's fence method (Correct answer)
Correct answer: Tukey's fence method
Tukey's fence method flags values beyond 1.5x IQR from Q1 or Q3 as outliers for removal or treatment.
Question 4: Under the EU AI Act, which risk category requires mandatory conformity assessments before deployment?
- Limited risk
- High risk (Correct answer)
- Minimal risk
- Prohibited risk
Correct answer: High risk
High-risk AI systems under the EU AI Act must undergo mandatory conformity assessments and meet strict requirements before market deployment.
Question 5: Which architecture is most commonly used for image-to-image tasks like semantic segmentation and medical image analysis?
- ResNet
- U-Net (encoder-decoder with skip connections) (Correct answer)
- VGG-16
- Inception v3
Correct answer: U-Net (encoder-decoder with skip connections)
U-Net uses a contracting encoder path and an expanding decoder path with skip connections that preserve spatial detail for precise pixel prediction.
Question 6: What does TF-IDF stand for in text feature extraction?
- Token Filter–Index Document Format
- Term Frequency–Inverse Document Frequency (Correct answer)
- Text Format–Inverse Data Frame
- Text Frequency–Inverse Document Function
Correct answer: Term Frequency–Inverse Document Frequency
TF-IDF weighs each word by how often it appears in a document (TF) penalized by how common it is across all documents (IDF).
Question 7: Which algorithm is commonly used in supervised learning?
- Autoencoders
- Linear Regression (Correct answer)
- PCA
- K-Means Clustering
Correct answer: Linear Regression
Linear Regression is a classic and widely used algorithm in supervised learning, specifically for regression tasks. It models the relationship between a dependent variable and one or more independent variables by fitting a linear equation to the observed data. This algorithm requires labeled data where the output variable is continuous, making it a prime example of a supervised approach.
Question 8: The concept of 'human-in-the-loop' is mandated for high-risk AI systems under the EU AI Act primarily to:
- Satisfy labor union requirements for workforce participation
- Reduce computational costs of inference
- Ensure meaningful human oversight and the ability to intervene or override AI decisions (Correct answer)
- Meet ISO 9001 quality management standards
Correct answer: Ensure meaningful human oversight and the ability to intervene or override AI decisions
The EU AI Act requires high-risk systems to allow human oversight so that operators can understand, monitor, and intervene in AI decisions to prevent harm.
Question 9: Which metric is typically used to evaluate classification algorithms?
- Mean Squared Error
- R-squared
- Silhouette Score
- Accuracy (Correct answer)
Correct answer: Accuracy
Accuracy is a widely used and intuitive metric for evaluating classification algorithms. It measures the proportion of correctly predicted instances (both true positives and true negatives) out of the total number of instances in the dataset. While other metrics like precision, recall, and F1-score provide more nuanced insights, accuracy offers a straightforward overall measure of a classifier's performance.
Question 10: When reporting model drift in production, which type of drift refers to changes in the input feature distribution?
- Covariate shift (Correct answer)
- Label drift
- Prior probability shift
- Concept drift
Correct answer: Covariate shift
Covariate shift occurs when the marginal distribution of input features P(X) changes while the conditional P(Y|X) remains stable.
Question 11: Which technique generates synthetic minority class samples to address class imbalance?
- Stratified sampling
- Bootstrapping
- Random undersampling
- SMOTE (Correct answer)
Correct answer: SMOTE
SMOTE creates new synthetic samples by interpolating between existing minority class examples rather than duplicating them.
Question 12: In MLflow, what is the purpose of the 'Model Registry' component?
- To schedule automated retraining pipelines
- To manage model versioning, staging, and production transitions (Correct answer)
- To store raw training datasets
- To visualize training metrics in real time
Correct answer: To manage model versioning, staging, and production transitions
MLflow Model Registry provides a centralized hub for managing the full lifecycle of ML models including versioning and stage transitions (Staging, Production, Archived).
Question 13: Which of the following BEST describes 'dual-use' risk in the context of machine learning research?
- Training a model on both labeled and unlabeled data simultaneously
- Research that can be applied for both beneficial and harmful purposes (Correct answer)
- Using two different ML frameworks for the same task
- Sharing a model across two separate business units
Correct answer: Research that can be applied for both beneficial and harmful purposes
Dual-use risk refers to technologies or research that can be intentionally or unintentionally repurposed to cause harm.
Question 14: In a typical CI/CD pipeline for ML (MLOps), what step immediately follows model training and precedes deployment?
- Feature engineering
- Infrastructure provisioning
- Model evaluation and validation (Correct answer)
- Data ingestion
Correct answer: Model evaluation and validation
After training, the model must pass evaluation gates (accuracy thresholds, fairness checks, performance benchmarks) before it is approved for deployment.
Question 15: Which evaluation metric for generative language models measures how well a probability model predicts a sample by computing the exponential of average negative log-likelihood?
- BLEU
- Perplexity (Correct answer)
- ROC-AUC
- F1 Score
Correct answer: Perplexity
Perplexity quantifies how surprised a language model is by new text; lower perplexity indicates better predictive performance.
Question 16: What is a machine learning data pipeline?
- A visualization tool for feature distributions
- A method for hyperparameter tuning
- An automated sequence of data processing steps from ingestion to model-ready format (Correct answer)
- A storage system for training datasets
Correct answer: An automated sequence of data processing steps from ingestion to model-ready format
A data pipeline automates and chains data collection, cleaning, transformation, and feature engineering steps for reproducibility.
Question 17: Which automated retraining trigger strategy is considered a best practice in production MLOps?
- Scheduled or event-driven retraining triggered by performance degradation thresholds or new data availability (Correct answer)
- Retraining only when a new model architecture is designed
- Manual retraining only on a fixed monthly schedule
- Retraining after every new user request to the API
Correct answer: Scheduled or event-driven retraining triggered by performance degradation thresholds or new data availability
Automated retraining pipelines that trigger on both schedules and performance thresholds keep models current without requiring manual intervention.
Question 18: Which method identifies predictive features by measuring how much each reduces impurity in a trained Random Forest?
- Variance Inflation Factor (VIF)
- Recursive Feature Elimination
- Feature importance from Random Forest (Correct answer)
- Correlation matrix analysis
Correct answer: Feature importance from Random Forest
Random Forest models compute feature importance scores based on average impurity reduction contributed by each feature across all trees.
Question 19: What is the primary purpose of the Expected Improvement (EI) acquisition function in Bayesian optimization?
- To ensemble predictions from multiple surrogates
- To balance exploration and exploitation by quantifying improvement probability over the current best (Correct answer)
- To directly minimize the surrogate model's prediction
- To sample uniformly from unexplored regions
Correct answer: To balance exploration and exploitation by quantifying improvement probability over the current best
EI computes the expected value of improvement over the current best observation, naturally trading off exploration (uncertainty) against exploitation (predicted value).
Question 20: What is the importance of continuing education in AML professional practice?
- It is only required for new professionals
- To justify higher fees
- To maintain current knowledge, meet certification requirements, and adapt to evolving industry standards (Correct answer)
- To fill time between client appointments
Correct answer: To maintain current knowledge, meet certification requirements, and adapt to evolving industry standards
Continuing education ensures professionals stay current with new developments, technologies, and best practices in their field while fulfilling certification renewal requirements and providing better service to clients.
Question 21: What is A/B testing in the context of ML model deployment?
- Testing models on two separate servers simultaneously
- Running automated unit tests on model training code
- Comparing two model versions by routing live traffic to each and measuring performance differences (Correct answer)
- Testing two different datasets against the same model
Correct answer: Comparing two model versions by routing live traffic to each and measuring performance differences
A/B testing routes a portion of real production traffic to a challenger model while the rest serves the incumbent, enabling data-driven model selection.
Question 22: What distinguishes a Restricted Boltzmann Machine (RBM) from a standard autoencoder as unsupervised representation learners?
- RBMs learn deterministic encodings; autoencoders learn probabilistic latent representations
- RBMs are generative probabilistic models; autoencoders are deterministic encoder-decoder architectures (Correct answer)
- RBMs use backpropagation; autoencoders use contrastive divergence
- RBMs require labeled data; autoencoders operate unsupervised
Correct answer: RBMs are generative probabilistic models; autoencoders are deterministic encoder-decoder architectures
RBMs are energy-based generative models trained with contrastive divergence that learn a probability distribution over inputs, while autoencoders deterministically encode inputs to reconstruct them.
Question 23: In AML practice, what is the purpose of a standard operating procedure (SOP)?
- To satisfy management preferences only
- To create unnecessary paperwork
- To restrict employee creativity
- To document step-by-step instructions for routine tasks to ensure consistency and quality (Correct answer)
Correct answer: To document step-by-step instructions for routine tasks to ensure consistency and quality
SOPs provide standardized, detailed instructions for routine operations, ensuring consistency, quality, efficiency, and safety regardless of which qualified individual performs the task.
Question 24: In TensorFlow Serving, which protocol is preferred for high-throughput, low-latency model inference in production?
- GraphQL
- WebSockets
- REST/HTTP
- gRPC (Correct answer)
Correct answer: gRPC
gRPC uses Protocol Buffers and HTTP/2 for binary serialization and multiplexing, providing lower latency and higher throughput than REST for TF Serving inference requests.
Question 25: Why is transparency important in AI systems?
- It helps users trust AI decisions (Correct answer)
- It prevents data sharing.
- It reduces computational costs.
- It hides proprietary algorithms.
Correct answer: It helps users trust AI decisions
Transparency in AI systems is crucial because it allows users to understand how and why an AI model arrived at a particular decision or prediction. When the decision-making process is clear and interpretable, it builds trust and confidence in the AI's outputs. This is especially important in sensitive applications where accountability and user acceptance are paramount.
Question 26: What fundamentally distinguishes SARSA from Q-learning?
- SARSA targets continuous action spaces while Q-learning targets discrete ones
- SARSA requires an environment model while Q-learning does not
- SARSA uses neural networks while Q-learning uses lookup tables
- SARSA is on-policy and updates using the actual next action taken; Q-learning is off-policy and updates using the greedy max-Q action (Correct answer)
Correct answer: SARSA is on-policy and updates using the actual next action taken; Q-learning is off-policy and updates using the greedy max-Q action
SARSA (on-policy) updates Q-values using the action actually taken by the behavior policy, while Q-learning (off-policy) always bootstraps from the greedy maximum Q-value action.
Question 27: What does professional liability insurance protect a AML practitioner against?
- Financial loss from claims of negligence, errors, or omissions in professional services (Correct answer)
- Natural disaster damage to the office
- Employee health expenses
- Theft of office equipment
Correct answer: Financial loss from claims of negligence, errors, or omissions in professional services
Professional liability insurance (errors and omissions coverage) protects practitioners from the financial consequences of claims alleging negligence, mistakes, or failure to perform professional duties, covering legal defense costs and settlements.
Question 28: What is the credit assignment problem in reinforcement learning?
- Distributing shared rewards fairly among agents in cooperative settings
- Determining how to allocate compute between training and inference
- Assigning numerical credit scores to competing reward function designs
- Identifying which past actions were responsible for a reward received much later in a sequence (Correct answer)
Correct answer: Identifying which past actions were responsible for a reward received much later in a sequence
The credit assignment problem is the difficulty of determining which earlier actions in a long trajectory caused a delayed reward, requiring techniques like eligibility traces or TD(λ) to propagate credit backward.
Question 29: What is the role of an ML model registry in the MLOps lifecycle?
- To serve as a central catalog for tracking, versioning, and managing trained model artifacts through the deployment lifecycle (Correct answer)
- To automate hyperparameter optimization across training runs
- To monitor live model predictions and alert on drift
- To store raw training datasets and feature CSVs
Correct answer: To serve as a central catalog for tracking, versioning, and managing trained model artifacts through the deployment lifecycle
A model registry provides a governed repository where teams can register, version, stage, and promote models from development to production.
Question 30: In the context of Neural Architecture Search (NAS), what does the 'weight sharing' strategy in one-shot NAS achieve?
- It trains multiple architectures sequentially to find the best
- It constrains the search to architectures with identical parameter counts
- It reuses pretrained weights from ImageNet for all candidate architectures
- It trains a single supernet where all candidate architectures share parameters, dramatically reducing search cost (Correct answer)
Correct answer: It trains a single supernet where all candidate architectures share parameters, dramatically reducing search cost
One-shot NAS trains a supernet once where all sub-architectures share parameters, then evaluates candidates by sampling from the supernet, reducing search cost from thousands of GPU-days to hours.
Question 31: What is semantic segmentation in computer vision?
- Ranking image regions by confidence score
- Generating a text caption for an image
- Assigning a class label to every pixel in an image (Correct answer)
- Detecting and localizing objects with bounding boxes
Correct answer: Assigning a class label to every pixel in an image
Semantic segmentation classifies each pixel of an image into a category, producing a pixel-wise class map of the entire scene.
Question 32: Which NLP technique converts words into dense vector representations that capture semantic relationships?
- Bag of Words
- Stemming
- TF-IDF
- Word2Vec / Word Embeddings (Correct answer)
Correct answer: Word2Vec / Word Embeddings
Word embeddings like Word2Vec map words to dense vectors where semantically similar words have high cosine similarity.
Question 33: What is the purpose of peer review in AML professional practice?
- To create competition between colleagues
- To determine who should be promoted
- To reduce workload through delegation
- To evaluate work quality through assessment by qualified colleagues and promote continuous improvement (Correct answer)
Correct answer: To evaluate work quality through assessment by qualified colleagues and promote continuous improvement
Peer review provides objective quality assessment by qualified professionals, identifying areas for improvement, validating practices, and promoting professional accountability and continuous quality enhancement.
Question 34: What should a AML professional do when they discover an ethical violation by a colleague?
- Report through appropriate organizational channels while documenting the observed behavior (Correct answer)
- Post about it on social media
- Ignore it to maintain workplace harmony
- Confront the colleague publicly
Correct answer: Report through appropriate organizational channels while documenting the observed behavior
Professionals have an ethical obligation to report violations through proper channels (supervisor, ethics committee, licensing board). Documentation of observations ensures accuracy and supports the investigation process.
Question 35: What is the purpose of model versioning in MLOps?
- To label different datasets used to train a model
- To increment model complexity with each training run
- To compress model weights for faster inference
- To track distinct model iterations with metadata to enable rollback and comparison (Correct answer)
Correct answer: To track distinct model iterations with metadata to enable rollback and comparison
Model versioning logs each trained model artifact with its configuration, metrics, and data lineage to support governance and rollback.
Question 36: What is the purpose of a train-validation-test split in model development?
- To tune hyperparameters and evaluate final model performance on unseen data (Correct answer)
- To increase the size of the training set
- To balance class distributions
- To perform feature engineering
Correct answer: To tune hyperparameters and evaluate final model performance on unseen data
The validation set enables hyperparameter tuning while the test set provides an unbiased final performance estimate.
Question 37: What is shadow deployment in machine learning operations?
- Deploying a model to a private server without public access
- Running a model only during off-peak hours to save resources
- Running a new model in parallel with the production model without using its predictions for actual decisions (Correct answer)
- Testing a model with synthetic data before live deployment
Correct answer: Running a new model in parallel with the production model without using its predictions for actual decisions
Shadow deployment routes live traffic to both the current and new model, logging new model outputs for evaluation without affecting end users.
Question 38: In AML practice, what is evidence-based practice?
- Using only the newest methods regardless of evidence
- Integrating the best available research evidence with professional expertise and client needs (Correct answer)
- Following only personal experience and intuition
- Doing whatever the client requests
Correct answer: Integrating the best available research evidence with professional expertise and client needs
Evidence-based practice combines rigorous research evidence, professional clinical expertise, and client/patient preferences and values to make informed decisions that optimize outcomes.
Question 39: What is the effect of using a very small batch size in stochastic gradient descent?
- Smoother gradient estimates with faster convergence
- Elimination of the need for learning rate scheduling
- Reduced memory usage with no effect on convergence quality
- Higher gradient noise that can escape local minima but slower wall-clock convergence (Correct answer)
Correct answer: Higher gradient noise that can escape local minima but slower wall-clock convergence
Small batch sizes produce noisy gradient estimates that can help escape sharp local minima but typically require more iterations and wall-clock time to converge.
Question 40: What is the purpose of subword tokenization methods like Byte-Pair Encoding (BPE) used in models like GPT?
- Convert text to binary representations for efficient storage
- Replace all punctuation tokens before model training
- Speed up inference by reducing vocabulary size to single characters
- Balance vocabulary size and out-of-vocabulary handling by splitting rare words into frequent subword units (Correct answer)
Correct answer: Balance vocabulary size and out-of-vocabulary handling by splitting rare words into frequent subword units
BPE iteratively merges frequent character pairs into subword units, allowing models to handle rare and unseen words by decomposing them.
Question 41: Which technique addresses the vanishing gradient problem by allowing gradients to flow directly through skip connections?
- Dropout regularization
- Residual connections (ResNets) (Correct answer)
- Batch normalization alone
- Weight decay
Correct answer: Residual connections (ResNets)
Residual connections let gradients bypass layers via identity shortcuts, preventing them from vanishing in very deep networks.
Question 42: In AML practice, what is the purpose of a standard operating procedure (SOP)?
- To create unnecessary paperwork
- To restrict employee creativity
- To satisfy management preferences only
- To document step-by-step instructions for routine tasks to ensure consistency and quality (Correct answer)
Correct answer: To document step-by-step instructions for routine tasks to ensure consistency and quality
SOPs provide standardized, detailed instructions for routine operations, ensuring consistency, quality, efficiency, and safety regardless of which qualified individual performs the task.
Question 43: What is the primary function of a model monitoring dashboard in production?
- To automate model retraining on a fixed schedule
- To display hyperparameter search results from training runs
- To track production metrics like prediction drift, latency, and data quality in real time (Correct answer)
- To visualize the model's internal architecture and weights
Correct answer: To track production metrics like prediction drift, latency, and data quality in real time
Monitoring dashboards give operations teams real-time visibility into model health, enabling rapid detection of and response to performance degradation.
Question 44: What does model quantization achieve during ML model deployment?
- It reduces model size and speeds up inference by using lower-precision numerical representations (Correct answer)
- It encrypts model weights for secure deployment in regulated environments
- It converts regression models into classification models automatically
- It increases model accuracy by adding more learnable parameters
Correct answer: It reduces model size and speeds up inference by using lower-precision numerical representations
Quantization converts 32-bit floating point weights to 8-bit integers, drastically reducing model size and inference latency with minimal accuracy loss.
Question 45: What is the primary advantage of hierarchical reinforcement learning (HRL) for complex tasks?
- It guarantees convergence to the global optimum in non-convex reward landscapes
- It decomposes complex long-horizon tasks into sub-goals, improving sample efficiency and enabling sub-behavior reuse (Correct answer)
- It enables parallel training across multiple processors with no communication overhead
- It eliminates the need for reward engineering by automatically deriving intrinsic reward signals
Correct answer: It decomposes complex long-horizon tasks into sub-goals, improving sample efficiency and enabling sub-behavior reuse
HRL breaks long-horizon tasks into hierarchies where high-level policies set sub-goals for lower-level policies, dramatically improving sample efficiency and allowing learned sub-behaviors to transfer across tasks.
Question 46: When a stakeholder misrepresents model capabilities in an external presentation, what is the most appropriate response?
- Publicly contradict the stakeholder during the presentation
- Ignore the misrepresentation to avoid conflict
- Allow the misrepresentation to stand to protect organizational reputation
- Privately correct the stakeholder with accurate information and offer to help prepare factually accurate materials for future communications (Correct answer)
Correct answer: Privately correct the stakeholder with accurate information and offer to help prepare factually accurate materials for future communications
Private correction paired with an offer to provide accurate materials addresses the immediate issue diplomatically while preventing future misrepresentations.
Question 47: Which technique is used to augment training images by randomly flipping, rotating, or cropping them?
- Data augmentation (Correct answer)
- Weight decay
- Dropout regularization
- Batch normalization
Correct answer: Data augmentation
Data augmentation artificially expands the training set with label-preserving transformations, reducing overfitting in vision models.
Question 48: What does professional liability insurance protect a AML practitioner against?
- Natural disaster damage to the office
- Employee health expenses
- Financial loss from claims of negligence, errors, or omissions in professional services (Correct answer)
- Theft of office equipment
Correct answer: Financial loss from claims of negligence, errors, or omissions in professional services
Professional liability insurance (errors and omissions coverage) protects practitioners from the financial consequences of claims alleging negligence, mistakes, or failure to perform professional duties, covering legal defense costs and settlements.
Question 49: In the context of ML risk management, 'operational risk' most directly refers to:
- The risk of choosing the wrong evaluation metric
- Overfitting caused by insufficient regularization
- Failures in people, processes, or systems supporting the model's operation (Correct answer)
- The chance that the chosen loss function is suboptimal
Correct answer: Failures in people, processes, or systems supporting the model's operation
Operational risk in ML covers failures stemming from inadequate internal processes, human error, system failures, or external events affecting the model's production pipeline.
Question 50: Which technique handles missing values by replacing them with the average of the column?
- Mean imputation (Correct answer)
- One-hot encoding
- Standardization
- Label encoding
Correct answer: Mean imputation
Mean imputation replaces missing values with the column's average, preserving the dataset's overall distribution.
Question 51: Which statistical test is recommended for comparing multiple ML models across multiple datasets to avoid inflated Type I error?
- Paired t-test with Bonferroni correction
- Friedman test followed by post-hoc Nemenyi test (Correct answer)
- ANOVA with Tukey's HSD
- Wilcoxon signed-rank test
Correct answer: Friedman test followed by post-hoc Nemenyi test
The Friedman test is a non-parametric rank-based test for multiple classifiers over multiple datasets, and the Nemenyi post-hoc test identifies which pairs differ significantly.
Question 52: What does CI/CD stand for in an MLOps pipeline?
- Continuous Improvement / Continuous Diagnostics
- Containerized Inference / Continuous Development
- Continuous Integration / Continuous Deployment (Correct answer)
- Central Intelligence / Cloud Deployment
Correct answer: Continuous Integration / Continuous Deployment
CI/CD automates testing and deployment pipelines so new model versions can be validated and released quickly and reliably.
Question 53: Which loss function is most appropriate for training a neural network on a multi-label classification task where each sample can belong to multiple classes?
- Binary cross-entropy applied independently per label (Correct answer)
- Categorical cross-entropy with softmax
- Mean squared error
- Hinge loss with multi-class margin
Correct answer: Binary cross-entropy applied independently per label
Binary cross-entropy treats each label independently with a sigmoid activation, allowing multiple classes to be simultaneously positive.
Question 54: Which feature engineering technique creates new features from products or powers of existing features?
- Polynomial feature expansion (Correct answer)
- Imputation
- Feature scaling
- Label encoding
Correct answer: Polynomial feature expansion
Polynomial feature expansion generates new features as products or powers of existing ones, enabling linear models to capture non-linear patterns.
Question 55: Which statement about t-SNE (t-Distributed Stochastic Neighbor Embedding) is TRUE?
- t-SNE is computationally efficient and scales well to millions of data points
- t-SNE is deterministic and always produces the same embedding
- t-SNE preserves global structure better than local structure
- t-SNE uses Student's t-distribution in low-dimensional space to alleviate the crowding problem (Correct answer)
Correct answer: t-SNE uses Student's t-distribution in low-dimensional space to alleviate the crowding problem
t-SNE uses a heavy-tailed t-distribution in the low-dimensional map, which allows moderately dissimilar points to be placed further apart and alleviates the crowding problem.
Question 56: Which dimensionality reduction technique projects data onto directions of maximum variance?
- UMAP
- LDA
- PCA (Correct answer)
- t-SNE
Correct answer: PCA
PCA (Principal Component Analysis) finds orthogonal axes of maximum variance to compress feature dimensions.
Question 57: What is the primary risk of using a static holdout test set for ongoing model quality monitoring over multiple retraining cycles?
- Overfitting to the validation set through hyperparameter tuning
- Test set contamination through repeated model selection on the same data (Correct answer)
- Increased inference cost per evaluation cycle
- Instability in precision-recall tradeoffs
Correct answer: Test set contamination through repeated model selection on the same data
Repeated selection of models based on the same test set causes the test set to implicitly influence model development, leaking information and inflating performance estimates.
Question 58: Which metric category is most critical when monitoring a deployed classification model in production?
- Training loss from the most recent training run
- Model file size on disk
- Prediction confidence distribution and accuracy on live data (Correct answer)
- Number of API calls per second
Correct answer: Prediction confidence distribution and accuracy on live data
Monitoring live prediction confidence and accuracy reveals model drift and performance degradation before it significantly impacts business outcomes.
Question 59: What is Inverse Reinforcement Learning (IRL) primarily designed to accomplish?
- Training agents through episodes played in reverse chronological order
- Inferring an unknown reward function from observed expert demonstrations (Correct answer)
- Learning policies that minimize rather than maximize cumulative reward
- Training agents to reverse or undo previously taken actions
Correct answer: Inferring an unknown reward function from observed expert demonstrations
IRL infers the underlying reward function from expert behavioral demonstrations, which is useful when the reward function is unknown, hard to define, or too complex to specify manually.
Question 60: In NLP, what does Named Entity Recognition (NER) identify in text?
- Duplicate sentences across a corpus
- Real-world entities such as people, organizations, and locations (Correct answer)
- The overall sentiment of a document
- The grammatical role of each word in a sentence
Correct answer: Real-world entities such as people, organizations, and locations
NER tags spans of text that refer to specific entities like person names, companies, dates, and geographic locations.
Question 61: What is the primary purpose of model explainability tools like SHAP in production ML systems?
- To automate feature engineering pipelines
- To compress model artifacts for faster deployment
- To speed up model inference at scale
- To explain individual predictions and ensure model transparency for stakeholders and auditors (Correct answer)
Correct answer: To explain individual predictions and ensure model transparency for stakeholders and auditors
SHAP attributes each feature's contribution to individual predictions, supporting transparency and regulatory compliance in production.
Question 62: What is the key distinction between model-based and model-free reinforcement learning?
- Model-free RL requires more labeled data than model-based approaches
- Model-based RL only works in discrete action spaces
- Model-based RL uses deep neural networks; model-free uses tabular methods only
- Model-based RL learns or uses an environment model for planning; model-free learns directly from interactions (Correct answer)
Correct answer: Model-based RL learns or uses an environment model for planning; model-free learns directly from interactions
Model-based RL builds or leverages an explicit model of environment dynamics for planning, while model-free RL learns policies or value functions directly from experience without modeling transitions.
Question 63: What does a log transformation primarily help with during feature engineering?
- Encoding categorical variables
- Reducing the skewness of right-skewed distributions (Correct answer)
- Handling missing data
- Normalizing binary features
Correct answer: Reducing the skewness of right-skewed distributions
Log transformation compresses large values and expands small ones, making right-skewed distributions more symmetric.
Question 64: Which tool is commonly used to containerize ML models for consistent deployment across environments?
- NumPy
- Jupyter Notebook
- Docker (Correct answer)
- Scikit-learn
Correct answer: Docker
Docker packages ML models with their dependencies into portable containers that run consistently across development and production environments.
Question 65: What is the primary benefit of feature selection in a machine learning pipeline?
- Reduces overfitting and training time (Correct answer)
- Adds synthetic training data
- Increases model complexity
- Improves data imputation accuracy
Correct answer: Reduces overfitting and training time
Feature selection removes irrelevant or redundant features, reducing overfitting risk and computational cost.
Question 66: An ML practitioner is hired to consult on a project but lacks expertise in the specific domain. According to professional ethics standards, the MOST appropriate response is to:
- Accept the project and learn on the job to avoid revenue loss
- Deliver results without mentioning the expertise gap
- Decline or disclose limitations and bring in a domain expert (Correct answer)
- Charge a lower fee to compensate for reduced quality
Correct answer: Decline or disclose limitations and bring in a domain expert
Professional ethics require practitioners to practice only within their competence and to be transparent about limitations.
Question 67: What is 'data versioning' in the context of ML tools like DVC (Data Version Control)?
- Encrypting sensitive training data before storage
- Compressing datasets to reduce storage costs
- Splitting datasets into versioned train/validation/test splits
- Tracking changes to large datasets and model artifacts using Git-like version control (Correct answer)
Correct answer: Tracking changes to large datasets and model artifacts using Git-like version control
DVC adds Git-like version control for large data files and ML artifacts by storing metadata in Git while the actual data lives in remote storage (S3, GCS, etc.).
Question 68: What is training-serving skew in machine learning deployment?
- The accuracy gap between training and test set performance
- The mismatch between model size and available hardware capacity
- The time lag between completing model training and completing deployment
- Discrepancies between how features are computed during offline training versus online model serving (Correct answer)
Correct answer: Discrepancies between how features are computed during offline training versus online model serving
Training-serving skew occurs when feature computation logic differs between offline training and online serving, causing unexpected prediction errors in production.
Question 69: What is model drift?
- The gradual improvement of a model over time
- A regularization technique for production models
- The movement of model weights during training
- The degradation of model performance as real-world data distribution changes from training data (Correct answer)
Correct answer: The degradation of model performance as real-world data distribution changes from training data
Model drift occurs when the statistical properties of input data or target relationships change post-deployment, reducing model accuracy.
Question 70: When publishing an ML research paper, a researcher finds that a competing team's unpublished work led them to the same solution. What is the ethically required action?
- Cite the competing team's work if it influenced the research (Correct answer)
- Contact the journal to claim priority based on submission date
- Publish independently since the competing work is unpublished
- Delay publication until the competing team publishes first
Correct answer: Cite the competing team's work if it influenced the research
Academic integrity requires attributing any work that influenced your research, regardless of its publication status.
Question 71: Which NLP preprocessing step reduces words to their base or root form by removing suffixes (e.g., 'running' → 'run')?
- Lemmatization
- POS tagging
- Dependency parsing
- Stemming (Correct answer)
Correct answer: Stemming
Stemming applies rule-based suffix stripping to reduce words to an approximate root, which may not be a valid dictionary word.
Question 72: What is MLOps in the context of advanced machine learning?
- The practice of combining ML development with operations to deploy and maintain models in production (Correct answer)
- A type of neural network architecture
- A dataset versioning tool
- A programming language for machine learning
Correct answer: The practice of combining ML development with operations to deploy and maintain models in production
MLOps applies DevOps principles to machine learning, automating training, deployment, monitoring, and retraining workflows.
Question 73: Equalized odds as a fairness criterion requires that a classifier satisfies:
- Equal precision across groups
- Equal positive prediction rates across groups
- Equal negative predictive value across groups
- Equal true positive and false positive rates across groups (Correct answer)
Correct answer: Equal true positive and false positive rates across groups
Equalized odds, defined by Hardt et al., requires both equal true positive rates (TPR) and equal false positive rates (FPR) across protected groups simultaneously.
Question 74: What is a REST API in the context of machine learning model deployment?
- A monitoring dashboard for deployed ML models
- A standardized interface that allows applications to send data to a model and receive predictions via HTTP (Correct answer)
- A database for storing model artifacts
- A version control system for datasets
Correct answer: A standardized interface that allows applications to send data to a model and receive predictions via HTTP
A REST API exposes model inference as HTTP endpoints so any downstream application can request predictions via standard web requests.
Question 75: Which stakeholder management approach is best when dealing with a highly technical stakeholder who wants to micromanage model architecture decisions?
- Block all communication from this stakeholder
- Exclude them from technical meetings entirely
- Agree with all their suggestions regardless of merit
- Establish a structured technical review process with defined decision rights and consultation points (Correct answer)
Correct answer: Establish a structured technical review process with defined decision rights and consultation points
A structured review process with defined decision rights respects technical expertise while maintaining project governance and team autonomy.
Question 76: In cohort analysis, what does retention rate measure?
- The fraction of users from an initial cohort who remain active in a later period (Correct answer)
- The churn rate of the bottom decile
- The average revenue per user cohort
- The percentage of new users acquired in a period
Correct answer: The fraction of users from an initial cohort who remain active in a later period
Retention rate tracks what fraction of users who joined in a given cohort period are still active N periods later.
Question 77: Which transformer-based model introduced bidirectional context for language understanding and set new NLP benchmarks in 2018?
- GPT-2
- BERT (Correct answer)
- ELMo
- Seq2Seq
Correct answer: BERT
BERT (Bidirectional Encoder Representations from Transformers) pre-trains on masked language modeling using full left and right context simultaneously.
Question 78: What role does a 'kill switch' or 'circuit breaker' play in ML model risk management?
- It terminates training early when validation loss plateaus
- It disables GPU acceleration during high-traffic periods
- It removes outlier training examples before each epoch
- It automatically halts or reverts model predictions when performance drops below a defined threshold (Correct answer)
Correct answer: It automatically halts or reverts model predictions when performance drops below a defined threshold
A kill switch or circuit breaker provides an automated safeguard that takes the model offline or reverts to a fallback rule-based system when real-time metrics signal unacceptable degradation.
Question 79: Why is k-fold cross-validation used when evaluating feature engineering choices?
- To increase the training dataset size
- To automatically select the optimal number of features
- To handle class imbalance
- To obtain a more reliable performance estimate across multiple data subsets (Correct answer)
Correct answer: To obtain a more reliable performance estimate across multiple data subsets
K-fold cross-validation rotates the validation set across k subsets, giving a more robust estimate than a single train-test split.
Question 80: What is the primary advantage of using Isolation Forest over LOF for anomaly detection?
- Isolation Forest is supervised and needs labeled anomalies
- Isolation Forest scales efficiently to high-dimensional, large datasets (Correct answer)
- Isolation Forest produces calibrated anomaly probabilities
- Isolation Forest requires no hyperparameters
Correct answer: Isolation Forest scales efficiently to high-dimensional, large datasets
Isolation Forest has linear time complexity and scales well to high dimensions and large datasets by isolating anomalies with random partitioning, unlike the quadratic-complexity LOF.
Question 81: What is the key difference between normalization and standardization?
- They are identical techniques with different names
- Normalization encodes categories; standardization scales numerics
- Normalization removes outliers; standardization does not
- Normalization scales values to [0,1]; standardization transforms data to zero mean and unit variance (Correct answer)
Correct answer: Normalization scales values to [0,1]; standardization transforms data to zero mean and unit variance
Normalization maps values to a fixed range like [0,1] while standardization transforms data to have mean 0 and standard deviation 1.
Question 82: What role does documentation play in AML regulatory compliance?
- It is only needed for tax purposes
- It creates unnecessary paperwork
- It is optional for experienced professionals
- It provides evidence of compliance, supports audits, and creates a defensible record of professional activities (Correct answer)
Correct answer: It provides evidence of compliance, supports audits, and creates a defensible record of professional activities
Documentation is the foundation of compliance, providing verifiable evidence that regulatory requirements have been met, supporting audit processes, and creating a defensible record if compliance is ever questioned.
Question 83: In Bayesian hyperparameter optimization, what probabilistic model is most commonly used to approximate the objective function?
- Neural network meta-learner
- Support vector regressor
- Gaussian Process (Correct answer)
- Random forest surrogate
Correct answer: Gaussian Process
Gaussian Processes provide a posterior distribution over the objective function, enabling uncertainty-aware acquisition functions like Expected Improvement.
Question 84: What is canary deployment in ML production systems?
- Testing new models in a completely isolated validation environment
- Using a lightweight fallback model when the primary model fails
- Gradually rolling out a new model to a small subset of users before a full release (Correct answer)
- Deploying a model only for internal QA testing
Correct answer: Gradually rolling out a new model to a small subset of users before a full release
Canary deployment incrementally increases traffic to a new model version, allowing real-world validation with minimal risk exposure.
Question 85: Which artifact BEST serves as the single source of truth for tracking ML experiment reproducibility across a team?
- A README file updated manually by each team member
- Slack messages with attached screenshots of training curves
- An MLflow or similar experiment tracking registry logging code version, data hash, params, and metrics (Correct answer)
- A shared Google Doc with hyperparameter tables
Correct answer: An MLflow or similar experiment tracking registry logging code version, data hash, params, and metrics
Experiment tracking platforms like MLflow capture code version, data lineage, hyperparameters, and metrics in a queryable, reproducible format.
Question 86: What is target encoding for categorical variables?
- Encoding the target variable as integers
- Replacing each category value with the mean of the target variable for that category (Correct answer)
- Encoding text features using TF-IDF
- Binarizing the target for classification tasks
Correct answer: Replacing each category value with the mean of the target variable for that category
Target encoding substitutes each category with the average target value for that category, directly capturing predictive power.
Question 87: A cross-functional team disagrees on model deployment timelines due to differing risk tolerances. What communication strategy should the ML lead use?
- Let each team implement their preferred timeline independently
- Force a technical decision without input from other teams
- Delay deployment indefinitely to avoid conflict
- Facilitate a structured risk assessment discussion using a shared framework to align all stakeholders (Correct answer)
Correct answer: Facilitate a structured risk assessment discussion using a shared framework to align all stakeholders
A shared risk assessment framework creates common ground and objective criteria that help cross-functional teams reach consensus on deployment timing.
Question 88: In NLP, what is tokenization?
- Removing stop words from a sentence
- Splitting raw text into individual units such as words or subwords (Correct answer)
- Converting words to their base grammatical form
- Mapping tokens to their embedding vectors
Correct answer: Splitting raw text into individual units such as words or subwords
Tokenization breaks raw text into discrete units (tokens) such as words, subwords, or characters that a model can process.
Question 89: What is 'tail risk' in the context of ML model predictions, and why is it critical for risk management?
- The risk of rare but extreme prediction errors that can cause disproportionately large negative outcomes (Correct answer)
- The risk that the last few layers of a neural network overfit
- The risk of slow inference for long-tail categorical features
- The risk that evaluation metrics are computed on the wrong data split
Correct answer: The risk of rare but extreme prediction errors that can cause disproportionately large negative outcomes
Tail risk refers to low-probability but high-severity errors at the extremes of the prediction distribution, which standard average-based metrics obscure but which can dominate real-world losses.
Question 90: What is transfer learning in the context of computer vision?
- Transferring labels from one dataset to another
- Using a pre-trained model's weights as a starting point for a new task (Correct answer)
- Training a model from scratch on a new dataset
- Copying model architecture without pre-trained weights
Correct answer: Using a pre-trained model's weights as a starting point for a new task
Transfer learning reuses the feature representations learned from a large dataset (e.g., ImageNet) to improve performance on a smaller target task.
Question 91: What is the primary function of 'gradient checkpointing' in deep learning?
- Saving model checkpoints when validation loss improves
- Trading compute for memory by recomputing activations during backpropagation instead of storing them (Correct answer)
- Checkpointing distributed training progress for fault tolerance
- Clipping gradients to prevent exploding gradient issues
Correct answer: Trading compute for memory by recomputing activations during backpropagation instead of storing them
Gradient checkpointing discards intermediate activations during the forward pass and recomputes them during backpropagation, reducing peak memory usage at the cost of ~33% more computation.
Question 92: What is 'concept drift' and why is it a risk management concern in deployed ML systems?
- Random variation in model outputs due to floating-point precision
- Instability in gradient descent caused by poor initialization
- Gradual memory leak in model serving infrastructure
- A change in the statistical relationship between inputs and the target variable over time (Correct answer)
Correct answer: A change in the statistical relationship between inputs and the target variable over time
Concept drift occurs when the underlying data-generating process changes so that previously learned input-output relationships no longer hold, degrading model reliability.
Question 93: What is the primary purpose of feature scaling in machine learning?
- To encode categorical variables
- To remove outliers from the dataset
- To ensure all features contribute equally to model training (Correct answer)
- To reduce the number of features
Correct answer: To ensure all features contribute equally to model training
Feature scaling normalizes feature ranges so that no single feature dominates model training due to its magnitude.
Question 94: Which encoding technique assigns integer values to categories that have a natural order?
- Ordinal encoding (Correct answer)
- Binary encoding
- Frequency encoding
- One-hot encoding
Correct answer: Ordinal encoding
Ordinal encoding maps ordered categories such as low/medium/high to integers that preserve their natural rank.
Question 95: Which NLP task involves assigning a label (positive, negative, neutral) to a piece of text based on its emotional tone?
- Coreference Resolution
- Named Entity Recognition
- Sentiment Analysis (Correct answer)
- Machine Translation
Correct answer: Sentiment Analysis
Sentiment analysis classifies text according to the opinion or emotion expressed, commonly as positive, negative, or neutral.
Question 96: What is the curse of dimensionality in machine learning?
- The difficulty of visualizing high-dimensional data
- The phenomenon where data becomes sparse as dimensions increase, degrading model performance (Correct answer)
- The problem of having too many training samples
- The computational cost of training deep neural networks
Correct answer: The phenomenon where data becomes sparse as dimensions increase, degrading model performance
As feature dimensions increase, data points become increasingly sparse, making distance-based algorithms less effective.
Question 97: What does the Variance Inflation Factor (VIF) measure in regression analysis?
- The importance of features in a regression model
- The variance of individual feature distributions
- The degree of multicollinearity among predictor variables (Correct answer)
- The inflation of model error as more features are added
Correct answer: The degree of multicollinearity among predictor variables
VIF quantifies how much the variance of a regression coefficient is inflated due to correlation with other predictor variables.
Question 98: Which cross-validation strategy is most appropriate when your dataset has significant temporal ordering?
- Shuffle-split CV
- Leave-one-out CV
- Time-series split (walk-forward) (Correct answer)
- Stratified k-fold
Correct answer: Time-series split (walk-forward)
Time-series split (walk-forward validation) prevents data leakage by always training on past data and validating on future data.
Advanced Machine Learning (AML) Certification
The AML certification validates advanced competency across the full machine learning lifecycle, covering core algorithms, feature engineering, model deployment, NLP, computer vision, and professional skills including AI ethics, governance, and stakeholder communication.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds