Data Science Exam — Questions and Answers
Question 1: A data scientist has just performed clustering on a dataset and wants to evaluate the quality of the resulting clusters without using any external labels. They choose to use the Davies-Bouldin Index. What would a low Davies-Bouldin Index score indicate?
- Good clustering, with low intra-cluster similarity and high inter-cluster similarity.
- Poor clustering, with high similarity between clusters.
- Good clustering, with clusters that are compact and well-separated from each other. (Correct answer)
- The algorithm failed to converge.
Correct answer: Good clustering, with clusters that are compact and well-separated from each other.
The Davies-Bouldin Index (DBI) evaluates clustering quality based on the ratio of within-cluster scatter to between-cluster separation. A lower DBI score indicates better clustering. This is because a low value signifies that the clusters are compact (low intra-cluster distance) and far apart from each other (high inter-cluster distance).
Question 2: Which Python library is most commonly used for constructing and training deep learning models?
- Pandas
- Scikit-learn
- Statsmodels
- PyTorch (Correct answer)
Correct answer: PyTorch
PyTorch (and TensorFlow/Keras) provide dynamic computation graphs and GPU acceleration essential for deep learning.
Question 3: What is the key difference between R² (coefficient of determination) and Adjusted R²?
- Adjusted R² penalizes for the number of predictors added (Correct answer)
- R² is only for linear models while Adjusted R² works for any model
- Adjusted R² is always higher than R²
- R² accounts for model complexity while Adjusted R² does not
Correct answer: Adjusted R² penalizes for the number of predictors added
Adjusted R² penalizes for adding irrelevant predictors, unlike R² which always increases or stays the same as features are added.
Question 4: What is the interquartile range (IQR)?
- The mean minus the median
- Q4 minus Q2
- Q1 minus the minimum value
- Q3 minus Q1 (Correct answer)
Correct answer: Q3 minus Q1
The IQR is Q3 − Q1, representing the middle 50% spread of data and useful for detecting outliers.
Question 5: Which visualization technique is specifically designed to compare distributions across multiple categorical groups simultaneously?
- Scatter plot
- Pie chart
- Grouped box plot (Correct answer)
- Line chart
Correct answer: Grouped box plot
A grouped box plot displays side-by-side box plots for each category, enabling direct comparison of median, spread, and outliers across groups.
Question 6: What is the purpose of the train-validation-test split in machine learning?
- To use training for learning, validation for hyperparameter tuning, and test for final unbiased evaluation (Correct answer)
- To reduce training time by using smaller datasets
- To prevent class imbalance from affecting model performance
- To ensure the model sees all data at some point during training
Correct answer: To use training for learning, validation for hyperparameter tuning, and test for final unbiased evaluation
Keeping a held-out test set ensures the final performance estimate is unbiased by any model selection decisions.
Question 7: Which NLP task involves identifying the grammatical role of each word in a sentence, such as noun or verb?
- Dependency parsing
- Named entity recognition
- Part-of-speech tagging (Correct answer)
- Semantic role labeling
Correct answer: Part-of-speech tagging
Part-of-speech tagging assigns grammatical categories (noun, verb, adjective, etc.) to each token in a sentence.
Question 8: Which technique is used in word2vec to predict surrounding context words given a target word?
- FastText
- Skip-gram (Correct answer)
- GloVe
- CBOW
Correct answer: Skip-gram
Skip-gram predicts surrounding context words given a center/target word, while CBOW does the opposite.
Question 9: What is the purpose of the 'CLS' token in BERT-style models?
- To indicate a masked token during pre-training
- To serve as an aggregate sequence representation for classification tasks (Correct answer)
- To mark the end of a sentence
- To separate two input sentences
Correct answer: To serve as an aggregate sequence representation for classification tasks
The [CLS] token is prepended to inputs, and its final hidden state is used as a pooled sentence-level representation for downstream classification.
Question 10: Which of the following feature selection methods is characterized by its use of a predictive model to evaluate the usefulness of a feature subset, but is computationally expensive due to its iterative nature?
- Dimensionality reduction
- Wrapper methods (Correct answer)
- Filter methods
- Embedded methods
Correct answer: Wrapper methods
Wrapper methods use a specific machine learning algorithm to evaluate the quality of a subset of features. They train and test the model with different feature subsets (e.g., through forward selection or backward elimination) to find the optimal combination, which makes them computationally intensive but often leads to better performance for the chosen model. [16, 17, 18]
Question 11: Which ensemble method trains base classifiers sequentially, with each model focusing more on previously misclassified examples?
- Bagging
- Stacking
- Boosting (Correct answer)
- Voting
Correct answer: Boosting
Boosting trains classifiers sequentially, reweighting training samples so subsequent models focus on the errors of prior ones.
Question 12: In shadow deployment for model validation in production, the new model:
- Uses synthetic data to simulate production conditions
- Runs in parallel receiving real traffic but its outputs are not served to users (Correct answer)
- Replaces the old model immediately after offline validation
- Is tested on a small random sample of live users
Correct answer: Runs in parallel receiving real traffic but its outputs are not served to users
Shadow deployment lets the new model process real traffic silently alongside the production model, enabling risk-free comparison of live performance.
Question 13: Which metric is most important when evaluating the performance of a big data storage system under write-heavy workloads?
- Write throughput measured in MB/s or records/second (Correct answer)
- Query response time for SELECT * operations
- Replication factor across data centers
- Number of supported SQL dialects
Correct answer: Write throughput measured in MB/s or records/second
Write throughput (records or bytes per second) directly measures how quickly the system can ingest data, which is the primary bottleneck in write-heavy workloads.
Question 14: What is Apache NiFi primarily designed for?
- Running distributed machine learning training jobs
- Providing a SQL interface to Kafka topics
- Automating and managing data flow between systems with a visual interface (Correct answer)
- Monitoring cluster resource utilization
Correct answer: Automating and managing data flow between systems with a visual interface
Apache NiFi provides a web-based graphical interface for designing, managing, and monitoring data flows between systems with built-in provenance tracking.
Question 15: What is the main purpose of a validation set as distinct from both training and test sets?
- To tune hyperparameters without biasing test evaluation (Correct answer)
- To provide more training data
- To detect data leakage
- To measure final model generalization
Correct answer: To tune hyperparameters without biasing test evaluation
The validation set is used for hyperparameter tuning so the test set remains unseen and provides an unbiased generalization estimate.
Question 16: When would you prefer a Decision Tree over Logistic Regression for classification?
- When you need probability calibration and linear decision boundaries
- When the dataset has very few features and a large number of samples
- When regularization of feature weights is the primary concern
- When the data has complex nonlinear feature interactions and interpretability is needed (Correct answer)
Correct answer: When the data has complex nonlinear feature interactions and interpretability is needed
Decision trees naturally capture nonlinear interactions without feature engineering, and their structure is human-interpretable via tree visualization.
Question 17: In Gaussian Mixture Models (GMM), what algorithm is used to estimate the model parameters?
- Singular Value Decomposition
- Gradient descent
- Expectation-Maximization (EM) (Correct answer)
- Principal Component Analysis
Correct answer: Expectation-Maximization (EM)
The EM algorithm alternates between assigning soft cluster membership probabilities (E-step) and updating Gaussian parameters (M-step) until convergence.
Question 18: What is the role of the key, query, and value matrices in scaled dot-product attention?
- They apply layer normalization before softmax
- They store the vocabulary embeddings for lookup
- They project inputs to compute attention scores and weighted value outputs (Correct answer)
- They encode positional information for each token
Correct answer: They project inputs to compute attention scores and weighted value outputs
Queries and keys are compared via dot product to produce attention weights, which then blend value vectors into the output.
Question 19: In sequential hypothesis testing, the sequential probability ratio test (SPRT) allows:
- Testing to continue indefinitely with no stopping rule
- Using a fixed sample size determined before data collection
- Controlling only the Type II error rate during data collection
- Early stopping when accumulated evidence strongly favors H₀ or H₁ (Correct answer)
Correct answer: Early stopping when accumulated evidence strongly favors H₀ or H₁
SPRT updates the likelihood ratio after each observation and stops sampling as soon as the ratio crosses predetermined boundaries for H₀ or H₁.
Question 20: What is transfer learning in deep learning?
- Transferring training data between datasets
- Converting a model from one framework to another
- Moving a trained model to a different hardware device
- Reusing a model pretrained on one task as the starting point for training on a related task (Correct answer)
Correct answer: Reusing a model pretrained on one task as the starting point for training on a related task
Transfer learning leverages features learned by a model on a large source task (e.g., ImageNet) and fine-tunes it on a smaller target task, dramatically reducing data and compute requirements.
Question 21: The concept of 'value alignment' in AI safety refers to:
- Aligning model weights to minimize loss functions
- Synchronizing model versions across distributed systems
- Ensuring AI systems pursue goals consistent with human values and intentions (Correct answer)
- Matching feature scales before training
Correct answer: Ensuring AI systems pursue goals consistent with human values and intentions
Value alignment ensures that as AI systems become more capable, their objectives remain consistent with what humans actually want and care about.
Question 22: What does a high AUC-ROC score but low AUC-PR score on the same model most likely indicate?
- The model is overfitting
- The dataset is highly imbalanced and the model struggles with the minority class (Correct answer)
- The features are poorly engineered
- The model is well-suited for the task
Correct answer: The dataset is highly imbalanced and the model struggles with the minority class
AUC-ROC can be misleadingly high on imbalanced data because TN dominates, while AUC-PR focuses on the minority class performance and reveals weaknesses.
Question 23: The principle of 'beneficence' in AI ethics requires that AI systems:
- Minimize the number of model parameters
- Actively promote the well-being of users and society (Correct answer)
- Operate within defined computational budgets
- Be open-source to enable peer review
Correct answer: Actively promote the well-being of users and society
Beneficence obligates AI developers to design systems that do good and improve human welfare, not merely avoid harm.
Question 24: What does the power of a hypothesis test measure?
- Probability of failing to reject H₀ when H₀ is false
- Probability of rejecting H₀ when H₀ is false (Correct answer)
- Probability of rejecting H₀ when H₀ is true
- Probability that the confidence interval contains the true parameter
Correct answer: Probability of rejecting H₀ when H₀ is false
Power = 1 − β, where β is the Type II error rate; it is the probability of correctly detecting a true effect.
Question 25: A loan approval model achieves equal accuracy across racial groups but still denies loans to minority applicants at a higher rate. Which fairness metric captures this disparity?
- Accuracy parity
- Demographic parity (statistical parity) (Correct answer)
- Calibration
- Individual fairness
Correct answer: Demographic parity (statistical parity)
Demographic parity measures whether the positive outcome rate is equal across groups, regardless of model accuracy within groups.
Question 26: A data analyst is using a box plot to examine the distribution of salaries for a specific job role. The plot reveals several data points located far beyond the whiskers of the box. What is the most likely interpretation of these points?
- The median salary
- Missing data values
- Potential outliers (Correct answer)
- The interquartile range (IQR)
Correct answer: Potential outliers
Box plots are a standard graphical EDA technique used to display the distribution of numerical data. The 'whiskers' typically extend to 1.5 times the interquartile range (IQR) from the first and third quartiles. Data points that fall outside of these whiskers are considered potential outliers that may require further investigation.
Question 27: Counterfactual explanations in AI ethics are used to:
- Generate synthetic data for fairness testing
- Show what minimal changes to input would alter a model's decision (Correct answer)
- Audit training pipelines for data leakage
- Detect adversarial attacks on models
Correct answer: Show what minimal changes to input would alter a model's decision
Counterfactual explanations tell users 'if X had been different, the outcome would have changed,' making decisions more actionable and understandable.
Question 28: Which practice helps ensure accountability when an AI system causes harm?
- Keeping model architecture proprietary
- Maintaining detailed audit logs of model decisions and data lineage (Correct answer)
- Reducing the number of stakeholders who can access the model
- Training models only on public datasets
Correct answer: Maintaining detailed audit logs of model decisions and data lineage
Audit logs create a traceable record of decisions and data flows, enabling post-hoc investigation and accountability when AI causes harm.
Question 29: A bank uses an AI model for loan approvals. It's discovered that the model denies loans to a disproportionately high number of applicants from a specific zip code, which is strongly correlated with a protected demographic group. Even though applicants' financial profiles are similar to approved applicants from other areas, their applications are rejected. What is the most likely ethical issue at play?
- Algorithmic bias (Correct answer)
- Insufficient training data volume
- Lack of model accuracy
- Poor feature engineering
Correct answer: Algorithmic bias
This scenario describes algorithmic bias, where an AI system produces systematically prejudiced outcomes due to erroneous assumptions in the machine learning process. In this case, the zip code is acting as a proxy for a protected attribute (e.g., race), causing discriminatory outcomes even if the protected attribute itself is not used.
Question 30: Which neural network architecture processes sequential data by maintaining a hidden state that captures information from previous time steps?
- Multilayer Perceptron
- CNN
- Autoencoder
- Recurrent Neural Network (RNN) (Correct answer)
Correct answer: Recurrent Neural Network (RNN)
RNNs have recurrent connections that pass the hidden state from one time step to the next, enabling them to model temporal dependencies in sequential data like text or time series.
Question 31: The Ljung-Box test applied to time series model residuals is used to test for what?
- Heteroscedasticity in the original series
- Cointegration between two time series
- Remaining autocorrelation in the residuals (Correct answer)
- Normality of the residual distribution
Correct answer: Remaining autocorrelation in the residuals
The Ljung-Box test checks whether the residuals from a fitted model show significant autocorrelation at multiple lags; significant autocorrelation means the model is inadequate.
Question 32: What capability does a SARIMA model add over a standard ARIMA model?
- Explicitly capturing repeating seasonal patterns (Correct answer)
- Learning long-range dependencies spanning decades
- Handling non-stationary series with structural breaks
- Modeling multiple correlated time series simultaneously
Correct answer: Explicitly capturing repeating seasonal patterns
SARIMA (Seasonal ARIMA) extends ARIMA by incorporating additional seasonal AR, differencing, and MA terms to model patterns that repeat at fixed seasonal intervals.
Question 33: What type of chart is best suited for showing the frequency distribution of a single continuous variable?
- Histogram (Correct answer)
- Scatter plot
- Bar chart
- Line chart
Correct answer: Histogram
A histogram divides continuous data into bins and displays their frequencies, making it ideal for showing the distribution of a single continuous variable.
Question 34: A retail company wants to test if there is a statistically significant difference in the average transaction value among customers using three different payment methods (Credit Card, Debit Card, Mobile Pay). Which statistical test is most appropriate for this analysis?
- Analysis of Variance (ANOVA) (Correct answer)
- Independent two-sample t-test
- Chi-squared test
- Paired t-test
Correct answer: Analysis of Variance (ANOVA)
ANOVA is used to compare the means of three or more independent groups. A t-test is only suitable for comparing the means of two groups. Using multiple t-tests would inflate the probability of a Type I error.
Question 35: What is the key difference between soft clustering and hard clustering?
- Soft clustering assigns probabilities of membership to multiple clusters; hard clustering assigns each point to exactly one cluster (Correct answer)
- Soft clustering is faster; hard clustering is more accurate
- Soft clustering uses distance metrics; hard clustering uses probability
- Soft clustering works on continuous data; hard clustering works on categorical data
Correct answer: Soft clustering assigns probabilities of membership to multiple clusters; hard clustering assigns each point to exactly one cluster
In soft (fuzzy) clustering like GMM, each point has a fractional membership probability across all clusters, unlike hard clustering where membership is binary.
Question 36: Which metric measures the proportion of actual positives correctly identified by a classifier?
- Recall (Correct answer)
- Accuracy
- F1 Score
- Precision
Correct answer: Recall
Recall (sensitivity) is TP / (TP + FN), measuring how well the model finds all true positives.
Question 37: What is the purpose of a Kafka offset?
- The byte position of a message within an Avro file
- The number of replicas for a Kafka topic
- A sequential identifier that tracks a consumer's read position within a partition (Correct answer)
- The lag between producer and broker in milliseconds
Correct answer: A sequential identifier that tracks a consumer's read position within a partition
An offset is a unique sequential number assigned to each message in a partition, allowing consumers to track and resume their read position.
Question 38: What is 'target leakage' in feature engineering?
- Normalizing the target variable before training
- Splitting the target variable into multiple outputs
- Including features that contain information not available at prediction time (Correct answer)
- Using the target variable to impute missing values in training data only
Correct answer: Including features that contain information not available at prediction time
Target leakage occurs when features used during training contain information that would not be available when the model makes real predictions.
Question 39: A model achieves 99% accuracy on a dataset where 99% of samples belong to one class. This is an example of:
- Excellent model performance
- Overfitting to training data
- The accuracy paradox (Correct answer)
- A well-calibrated model
Correct answer: The accuracy paradox
The accuracy paradox occurs when high accuracy is misleading because a naive baseline (predicting the majority class always) achieves the same score.
Question 40: When using k-Nearest Neighbors for classification, what is the effect of choosing a very large k?
- The model can only handle binary classification
- Computation time decreases
- The decision boundary becomes smoother and simpler (Correct answer)
- The model becomes more sensitive to noise
Correct answer: The decision boundary becomes smoother and simpler
Large k values smooth out the decision boundary by averaging over more neighbors, reducing variance but potentially increasing bias.
Question 41: In the context of maximum likelihood estimation, the score function is:
- The first derivative of the log-likelihood with respect to θ (Correct answer)
- The second derivative of the log-likelihood with respect to θ
- The ratio of the likelihood to its maximum value
- The expected value of the log-likelihood under the true parameter
Correct answer: The first derivative of the log-likelihood with respect to θ
The score function S(θ) = ∂ℓ/∂θ equals zero at the MLE and has expectation zero under the true parameter.
Question 42: What does Spearman's rank correlation measure that Pearson's correlation does not?
- Monotonic relationships, including nonlinear ones, between two variables (Correct answer)
- The linear relationship between two variables
- The covariance between two variables scaled by their variances
- The proportion of variance explained by a linear fit
Correct answer: Monotonic relationships, including nonlinear ones, between two variables
Spearman's correlation computes Pearson's r on the ranks of the data, capturing any monotonic relationship (not just linear) and being robust to outliers.
Question 43: In a Kappa Architecture, what replaces the batch layer found in Lambda Architecture?
- A separate Spark cluster running nightly batch jobs
- A dedicated OLAP cube
- Reprocessing historical data through the same stream processing pipeline with a replay mechanism (Correct answer)
- A relational database for serving historical queries
Correct answer: Reprocessing historical data through the same stream processing pipeline with a replay mechanism
Kappa Architecture eliminates the separate batch layer by reprocessing historical data by replaying it through the unified stream processing pipeline, simplifying operations.
Question 44: Which of the following is an example of an extractive summarization approach?
- Generating new sentences that paraphrase the source document
- Selecting and concatenating important sentences directly from the source (Correct answer)
- Training a classifier to rank document quality
- Using a language model to rewrite the document in a shorter form
Correct answer: Selecting and concatenating important sentences directly from the source
Extractive summarization selects actual sentences or phrases from the original document rather than generating new text.
Question 45: When would you prefer forward feature selection over backward feature elimination?
- When the model does not support partial feature sets
- When you have very few features and need to remove some
- When all features are continuous
- When the number of features is large relative to samples, making fitting a full model impractical (Correct answer)
Correct answer: When the number of features is large relative to samples, making fitting a full model impractical
Forward selection starts with no features and adds one at a time, avoiding the need to fit a model on the full high-dimensional feature set.
Data Science Exam
The Data Science Examination (DSE) evaluates proficiency in computer science, mathematics, and statistics for data science professionals, administered by Pearson VUE.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds