Data Science Exam — Questions and Answers
Question 1: The principle of 'data minimization' in AI ethics states that:
- Datasets must be compressed before storage
- Feature counts should not exceed 100 per model
- Only data strictly necessary for the stated purpose should be collected (Correct answer)
- Models should use as few parameters as possible
Correct answer: Only data strictly necessary for the stated purpose should be collected
Data minimization reduces privacy risks by ensuring organizations collect only the data they genuinely need, limiting exposure in case of a breach.
Question 2: In the context of maximum likelihood estimation, the score function is:
- The expected value of the log-likelihood under the true parameter
- The second derivative of the log-likelihood with respect to θ
- The ratio of the likelihood to its maximum value
- The first derivative of the log-likelihood with respect to θ (Correct answer)
Correct answer: The first derivative of the log-likelihood with respect to θ
The score function S(θ) = ∂ℓ/∂θ equals zero at the MLE and has expectation zero under the true parameter.
Question 3: Select the option that centers around identifying unknown properties in the data from the list below.
- Big data
- Data wrangling
- Data mining (Correct answer)
- Machine learning
Correct answer: Data mining
Data mining is the process of discovering patterns, anomalies, and correlations within large datasets to predict outcomes. Its core objective is to extract previously unknown, useful information and insights from data, often using a combination of statistical methods, machine learning, and database systems. This aligns perfectly with identifying 'unknown properties' or hidden structures within the data.
Question 4: A one-sample chi-squared test is used to test whether observed category frequencies match:
- Two sample variances are equal
- Expected frequencies from a theoretical distribution (Correct answer)
- A specified mean vector
- Regression coefficients are jointly zero
Correct answer: Expected frequencies from a theoretical distribution
The chi-squared goodness-of-fit test compares observed counts to expected counts derived from a hypothesized distribution.
Question 5: Which metric is commonly used to evaluate named entity recognition (NER) performance?
- Entity-level F1 score (Correct answer)
- ROUGE-L
- Perplexity
- BLEU score
Correct answer: Entity-level F1 score
NER is evaluated with entity-level F1, which requires both the entity boundary and entity type to match for a prediction to count as correct.
Question 6: A bank uses an AI model for loan approvals. It's discovered that the model denies loans to a disproportionately high number of applicants from a specific zip code, which is strongly correlated with a protected demographic group. Even though applicants' financial profiles are similar to approved applicants from other areas, their applications are rejected. What is the most likely ethical issue at play?
- Lack of model accuracy
- Poor feature engineering
- Algorithmic bias (Correct answer)
- Insufficient training data volume
Correct answer: Algorithmic bias
This scenario describes algorithmic bias, where an AI system produces systematically prejudiced outcomes due to erroneous assumptions in the machine learning process. In this case, the zip code is acting as a proxy for a protected attribute (e.g., race), causing discriminatory outcomes even if the protected attribute itself is not used.
Question 7: In a confusion matrix for binary classification, what does a False Negative (FN) represent?
- Model predicted positive; actual class is positive
- Model predicted negative; actual class is negative
- Model predicted positive; actual class is negative
- Model predicted negative; actual class is positive (Correct answer)
Correct answer: Model predicted negative; actual class is positive
A False Negative occurs when the model predicts the negative class but the true label is positive, meaning a real positive was missed.
Question 8: When deploying an AI system in a criminal sentencing context, which concern is MOST ethically critical?
- Potential for the model to perpetuate systemic racial bias in sentencing (Correct answer)
- Model interpretability for software engineers
- The licensing cost of the ML framework used
- Inference latency of the model
Correct answer: Potential for the model to perpetuate systemic racial bias in sentencing
Criminal sentencing has irreversible consequences on individuals' lives, making fairness and bias the paramount ethical concern.
Question 9: In spectral clustering, what is the role of the Laplacian matrix?
- Normalizes feature values
- Computes pairwise Euclidean distances
- Selects the number of clusters automatically
- Encodes graph connectivity for eigendecomposition (Correct answer)
Correct answer: Encodes graph connectivity for eigendecomposition
The graph Laplacian captures the connectivity structure of the data, and its eigenvectors reveal cluster membership in the embedded space.
Question 10: Which technique adjusts a classifier's predicted probabilities to be better calibrated without retraining it?
- Grid search
- SMOTE
- Principal Component Analysis
- Platt scaling (Correct answer)
Correct answer: Platt scaling
Platt scaling fits a logistic regression on the model's raw outputs to transform them into calibrated probabilities.
Question 11: Which optimization algorithm combines momentum with adaptive per-parameter learning rates and is widely used as a default optimizer for deep learning?
- SGD
- Adagrad
- RMSProp
- Adam (Correct answer)
Correct answer: Adam
Adam (Adaptive Moment Estimation) maintains both first-moment (momentum) and second-moment (per-parameter adaptive) estimates, making it robust and fast-converging for most deep learning tasks.
Question 12: What is the main advantage of using a KDE (Kernel Density Estimate) over a histogram for EDA?
- It is computationally faster to compute
- It shows exact counts of observations in each interval
- It produces a smooth continuous density curve that is less sensitive to bin width choices (Correct answer)
- It requires no assumption about the underlying distribution
Correct answer: It produces a smooth continuous density curve that is less sensitive to bin width choices
KDE creates a smooth density estimate that avoids the arbitrary bin-width problem of histograms, giving a clearer view of the distribution shape.
Question 13: What does the K imply algorithm's K stand for?
- Number of clusters (Correct answer)
- Number of attributes
- Number of iterations
- Number of data
Correct answer: Number of clusters
In the K-means clustering algorithm, the 'K' explicitly stands for the number of clusters that the algorithm will attempt to identify within the dataset. The user must specify this 'K' value beforehand, guiding the algorithm to partition the data into that many distinct groups. The algorithm then iteratively assigns data points to the nearest cluster centroid and updates the centroids until convergence.
Question 14: Which of the following is an example of categorical data?
- Number of sales per day
- Annual revenue in dollars
- Customer satisfaction level (Low/Medium/High) (Correct answer)
- Temperature in Celsius
Correct answer: Customer satisfaction level (Low/Medium/High)
Categorical data represents discrete groups or categories rather than numerical measurements.
Question 15: What is the primary purpose of a Q-Q (quantile-quantile) plot in data analysis?
- Visualizing cluster assignments
- Assessing whether a dataset follows a specific theoretical distribution (Correct answer)
- Plotting two categorical variables
- Comparing two time series
Correct answer: Assessing whether a dataset follows a specific theoretical distribution
A Q-Q plot compares the quantiles of sample data against the quantiles of a reference distribution to assess how well the data fits that distribution.
Question 16: What is the main advantage of subword tokenization methods like Byte-Pair Encoding (BPE) over word-level tokenization?
- Handles out-of-vocabulary words more gracefully (Correct answer)
- Faster training speed
- Produces shorter sequences
- Requires no preprocessing
Correct answer: Handles out-of-vocabulary words more gracefully
BPE breaks rare and unknown words into subword units, reducing out-of-vocabulary issues while keeping common words intact.
Question 17: Which of the following best describes the bias-variance tradeoff?
- Both bias and variance decrease as model complexity increases
- High bias causes overfitting; high variance causes underfitting
- Bias and variance are independent of model complexity
- High bias causes underfitting; high variance causes overfitting (Correct answer)
Correct answer: High bias causes underfitting; high variance causes overfitting
Simple models have high bias (underfitting) while complex models have high variance (overfitting).
Question 18: Counterfactual explanations in AI ethics are used to:
- Detect adversarial attacks on models
- Audit training pipelines for data leakage
- Show what minimal changes to input would alter a model's decision (Correct answer)
- Generate synthetic data for fairness testing
Correct answer: Show what minimal changes to input would alter a model's decision
Counterfactual explanations tell users 'if X had been different, the outcome would have changed,' making decisions more actionable and understandable.
Question 19: How many groups in total can data be characterized?
- 4
- 2 (Correct answer)
- 3
- 1
Correct answer: 2
Data can broadly be characterized into two fundamental groups: qualitative (categorical) and quantitative (numerical). Qualitative data describes qualities or characteristics that cannot be measured numerically, while quantitative data consists of numerical values that can be measured or counted. These two types form the basis for all data collection, analysis, and modeling approaches.
Question 20: In dependency parsing, what does a dependency arc represent?
- A phrase boundary between two constituents
- A directed grammatical relationship between a head word and a dependent word (Correct answer)
- A semantic role assigned to a verb argument
- A co-reference link between two noun phrases
Correct answer: A directed grammatical relationship between a head word and a dependent word
Dependency arcs connect a head word to its grammatical dependent, labeled with the type of syntactic relation (e.g., subject, object).
Question 21: Which of the following best describes 'polynomial feature expansion'?
- Generating new features as powers and cross-products of original features up to a specified degree (Correct answer)
- Encoding each category as a polynomial function
- Applying a polynomial activation function to feature values
- Reducing features using polynomial regression
Correct answer: Generating new features as powers and cross-products of original features up to a specified degree
Polynomial expansion creates new features like x², x³, and x₁·x₂ to allow linear models to fit non-linear relationships.
Question 22: Which of the following is an example of 'automation bias' in an AI-assisted medical diagnostic system?
- The model flags incorrect lab values due to data entry errors
- The model learns from unbalanced class distributions
- Clinicians over-rely on the AI recommendation and skip their own assessment (Correct answer)
- The system automates billing codes incorrectly
Correct answer: Clinicians over-rely on the AI recommendation and skip their own assessment
Automation bias is the tendency for humans to defer to automated systems even when their own judgment would yield a better result.
Question 23: What does PCA (Principal Component Analysis) primarily accomplish?
- Removes duplicate records from a dataset
- Clusters data points into groups
- Reduces dimensionality by projecting data onto directions of maximum variance (Correct answer)
- Classifies data into predefined categories
Correct answer: Reduces dimensionality by projecting data onto directions of maximum variance
PCA transforms features into a smaller set of uncorrelated principal components that capture the most variance in the data.
Question 24: The principle of 'beneficence' in AI ethics requires that AI systems:
- Actively promote the well-being of users and society (Correct answer)
- Minimize the number of model parameters
- Be open-source to enable peer review
- Operate within defined computational budgets
Correct answer: Actively promote the well-being of users and society
Beneficence obligates AI developers to design systems that do good and improve human welfare, not merely avoid harm.
Question 25: Which chart type is most appropriate for comparing proportions that sum to a whole?
- Heatmap
- Histogram
- Scatter plot
- Pie chart (Correct answer)
Correct answer: Pie chart
A pie chart divides a circle into slices proportional to each category's share of the total, making part-to-whole comparisons intuitive.
Question 26: What is the purpose of one-hot encoding in machine learning preprocessing?
- To normalize numerical features to a 0-1 range
- To impute missing values using the most frequent category
- To reduce the number of categories by merging rare levels
- To convert categorical variables into binary indicator columns for use in algorithms requiring numerical input (Correct answer)
Correct answer: To convert categorical variables into binary indicator columns for use in algorithms requiring numerical input
One-hot encoding represents each category as a separate binary column, allowing algorithms that require numerical input to use categorical variables.
Question 27: When using Fuzzy C-Means clustering, what does the fuzziness parameter 'm' control?
- The distance metric used
- The number of clusters
- The maximum number of iterations
- The degree of overlap between cluster memberships (Correct answer)
Correct answer: The degree of overlap between cluster memberships
Higher values of m (m > 1) increase fuzziness so memberships are spread more evenly across clusters; as m → 1, Fuzzy C-Means approaches hard K-Means partitioning.
Question 28: What is 'data leakage' in the context of model validation?
- Information from the test set influencing model training or evaluation (Correct answer)
- Data accidentally deleted during preprocessing
- Missing values in the training set
- Overfitting due to too many features
Correct answer: Information from the test set influencing model training or evaluation
Data leakage occurs when information from outside the training boundary (e.g., test labels or future data) influences the model, causing overly optimistic results.
Question 29: A model trained on hospital A data is tested on hospital B data and performs much worse. This primarily illustrates:
- Distribution shift / covariate shift (Correct answer)
- Model underfitting
- Overfitting to training data
- Label noise
Correct answer: Distribution shift / covariate shift
Distribution shift occurs when the statistical properties of the test environment differ from training, leading to performance degradation.
Question 30: What is a key difference between Exploratory Data Analysis (EDA) and Confirmatory Data Analysis (CDA), such as formal hypothesis testing?
- EDA focuses on data cleaning, while CDA focuses on model building.
- EDA is only performed on small datasets, while CDA is used for big data.
- EDA is an open-ended process of generating hypotheses, while CDA aims to rigorously test pre-specified hypotheses. (Correct answer)
- EDA uses graphical techniques, while CDA exclusively uses statistical calculations.
Correct answer: EDA is an open-ended process of generating hypotheses, while CDA aims to rigorously test pre-specified hypotheses.
EDA is an exploratory, open-ended approach where the goal is to discover patterns, generate questions, and form hypotheses from the data without preconceived notions. In contrast, Confirmatory Data Analysis (which includes hypothesis testing) is a more rigid process focused on evaluating the evidence for or against specific hypotheses that were formulated beforehand.
Question 31: When using k-Nearest Neighbors for classification, what is the effect of choosing a very large k?
- The model can only handle binary classification
- The decision boundary becomes smoother and simpler (Correct answer)
- Computation time decreases
- The model becomes more sensitive to noise
Correct answer: The decision boundary becomes smoother and simpler
Large k values smooth out the decision boundary by averaging over more neighbors, reducing variance but potentially increasing bias.
Question 32: When should you prefer a log scale on a histogram's x-axis?
- When the sample size is small
- When the distribution is perfectly normal
- When the variable has negative values
- When the variable spans several orders of magnitude (Correct answer)
Correct answer: When the variable spans several orders of magnitude
A log scale is ideal for variables spanning several orders of magnitude (e.g., income, population) so that all ranges are visually represented proportionally.
Question 33: When evaluating fairness in a recidivism prediction model, which of the following is a documented real-world criticism of the COMPAS system?
- It was trained on synthetic data without real criminal records
- It used too few features to make predictions
- It produced higher false positive rates for Black defendants than white defendants (Correct answer)
- It ignored socioeconomic features entirely
Correct answer: It produced higher false positive rates for Black defendants than white defendants
ProPublica's 2016 analysis found COMPAS had significantly higher false positive rates for Black defendants, flagging them as high-risk when they did not reoffend.
Question 34: A company uses a facial recognition model that has a 1% error rate for light-skinned males but a 35% error rate for dark-skinned females. This performance gap is primarily a concern under which ethical principle?
- Intellectual property
- Scalability
- Efficiency
- Fairness and non-discrimination (Correct answer)
Correct answer: Fairness and non-discrimination
Disparate error rates across demographic groups violate fairness and non-discrimination principles central to AI ethics.
Question 35: Which file format is most commonly used to exchange tabular data between systems?
- Parquet
- JSON
- XML
- CSV (Correct answer)
Correct answer: CSV
CSV (Comma-Separated Values) is the most universally supported format for simple tabular data exchange.
Question 36: Which of the following is NOT a common data cleaning task?
- Fixing inconsistent formatting
- Training a neural network (Correct answer)
- Removing duplicate rows
- Handling missing values
Correct answer: Training a neural network
Training a neural network is a modeling step, not a data cleaning task like imputation or deduplication.
Question 37: Which loss function is typically used for binary classification problems in neural networks?
- Binary cross-entropy (Correct answer)
- Hinge loss
- Categorical cross-entropy
- Mean squared error
Correct answer: Binary cross-entropy
Binary cross-entropy measures the dissimilarity between the predicted probability and the true binary label, penalizing confident wrong predictions most heavily.
Question 38: Which EDA technique would you use to detect multicollinearity among predictor variables before modeling?
- Correlation matrix or variance inflation factors (Correct answer)
- Scatter plot of target vs. one predictor
- Histogram of the target variable
- Box plot of each predictor
Correct answer: Correlation matrix or variance inflation factors
A correlation matrix reveals linear dependencies between predictors, and VIF quantifies how much one predictor's variance is explained by others.
Question 39: Which ensemble method trains base classifiers sequentially, with each model focusing more on previously misclassified examples?
- Boosting (Correct answer)
- Bagging
- Voting
- Stacking
Correct answer: Boosting
Boosting trains classifiers sequentially, reweighting training samples so subsequent models focus on the errors of prior ones.
Question 40: What language is utilized in the field of data science?
- Ruby (Correct answer)
- Java
- c++
- R
Correct answer: Ruby
While Python and R are dominant in data science, Ruby is a general-purpose language that can also be utilized. It offers capabilities for data manipulation, scripting, and web development, which can be relevant in data-related projects, especially when integrating with web applications. Although its ecosystem for advanced statistical modeling and machine learning is less extensive than Python or R, Ruby's flexibility allows for its use in various data science tasks.
Question 41: Which of the following is a sign of high variance (overfitting) when examining learning curves?
- Training error is low but validation error is much higher (Correct answer)
- Both training and validation error are high
- Training error equals validation error and both are high
- Validation error decreases then plateaus
Correct answer: Training error is low but validation error is much higher
A large gap between low training error and high validation error is the classic signature of overfitting (high variance).
Question 42: What is the main difference between divisive and agglomerative hierarchical clustering?
- Divisive requires specifying k upfront; agglomerative does not
- Divisive clustering uses Euclidean distance; agglomerative uses cosine similarity
- Divisive starts with one cluster and splits; agglomerative starts with n clusters and merges (Correct answer)
- Divisive is faster; agglomerative produces better results
Correct answer: Divisive starts with one cluster and splits; agglomerative starts with n clusters and merges
Agglomerative (bottom-up) begins with each point as its own cluster and merges, while divisive (top-down) begins with all points in one cluster and recursively splits.
Question 43: Which of the following tasks is best described as a seq2seq problem?
- Sentiment classification
- Part-of-speech tagging
- Text summarization (Correct answer)
- Word sense disambiguation
Correct answer: Text summarization
Text summarization takes a long input sequence and generates a shorter output sequence, matching the encoder-decoder seq2seq paradigm.
Question 44: What does 'partition pruning' mean in the context of big data query optimization?
- Skipping the scan of partitions that cannot contain data matching the query's filter conditions (Correct answer)
- Splitting large partitions into smaller ones for better parallelism
- Deleting old partitions based on a retention policy
- Removing corrupted partitions from HDFS
Correct answer: Skipping the scan of partitions that cannot contain data matching the query's filter conditions
Partition pruning allows the query engine to read only the partitions relevant to the filter predicate, dramatically reducing I/O on large partitioned datasets.
Question 45: In sentiment analysis, what is the challenge of handling negation in sentences like 'not good'?
- Neural models cannot process negation at all
- Negation is always removed during stopword filtering
- Negation only affects named entities
- Bag-of-words models miss the interaction between 'not' and 'good' (Correct answer)
Correct answer: Bag-of-words models miss the interaction between 'not' and 'good'
Simple bag-of-words models treat words independently, so 'not' and 'good' are scored separately rather than as an inverted sentiment unit.
Data Science Exam
The Data Science Examination (DSE) evaluates proficiency in computer science, mathematics, and statistics for data science professionals, administered by Pearson VUE.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds