Data Science Certification Exam — Questions and Answers
Question 1: Which technique is used to reduce the number of features in a dataset while preserving most of the variance?
- Decision Tree
- K-Means Clustering
- Linear Regression
- Principal Component Analysis (Correct answer)
Correct answer: Principal Component Analysis
PCA reduces dimensionality by transforming features into principal components that capture the maximum variance in the data.
Question 2: In pandas, which method removes rows containing any missing values from a DataFrame?
- df.isnull()
- df.dropna() (Correct answer)
- df.drop_duplicates()
- df.fillna()
Correct answer: df.dropna()
dropna() drops rows (or columns) that contain NaN values by default.
Question 3: What type of machine learning algorithm is K-Means?
- Semi-supervised regression
- Unsupervised clustering (Correct answer)
- Reinforcement learning
- Supervised classification
Correct answer: Unsupervised clustering
K-Means is an unsupervised clustering algorithm that partitions data into K distinct groups based on feature similarity without using labeled outcomes.
Question 4: In the context of machine learning, what is the primary purpose of a validation set?
- To tune hyperparameters and select the best model configuration (Correct answer)
- To train the model's parameters
- To evaluate the final model performance on unseen data
- To preprocess and clean the raw input data
Correct answer: To tune hyperparameters and select the best model configuration
A validation set is used during development to tune hyperparameters and compare different model configurations. The test set (not the validation set) is reserved for final unbiased evaluation.
Question 5: Which algorithm builds an ensemble of trees sequentially, each correcting the errors of the previous?
- Gradient boosting (Correct answer)
- Random forest
- K-nearest neighbors
- Naive Bayes
Correct answer: Gradient boosting
Gradient boosting adds trees sequentially to minimize the residual errors of prior trees.
Question 6: When using SHAP values for feature analysis, they explain:
- The number of categories
- Only the global dataset mean
- The training time per feature
- Each feature's contribution to an individual prediction (Correct answer)
Correct answer: Each feature's contribution to an individual prediction
SHAP values attribute a prediction's deviation from baseline to each contributing feature.
Question 7: Which pandas method removes duplicate rows?
- df.drop_duplicates() (Correct answer)
- df.unique()
- df.dropna()
- df.distinct()
Correct answer: df.drop_duplicates()
drop_duplicates() removes repeated rows from a DataFrame.
Question 8: Which statement about correlation and causation is correct?
- Correlation always implies causation
- Correlation does not imply causation (Correct answer)
- A high correlation proves one variable causes the other
- Causation requires zero correlation
Correct answer: Correlation does not imply causation
Two variables can correlate due to confounders or coincidence without any causal link.
Question 9: The interquartile range (IQR) measures:
- The difference between the maximum and minimum
- The spread of the middle 50% of the data (Correct answer)
- The average distance from the mean
- The most frequent value
Correct answer: The spread of the middle 50% of the data
The IQR is the distance between the 25th and 75th percentiles, capturing the central half of the data.
Question 10: A chi-square test of independence is used to assess:
- Difference between two means
- Linear correlation of two numeric variables
- Association between two categorical variables (Correct answer)
- Equality of variances
Correct answer: Association between two categorical variables
The chi-square test of independence evaluates whether two categorical variables are associated.
Question 11: What is the main advantage of using mutual information over Pearson correlation for feature selection?
- It captures non-linear relationships between variables (Correct answer)
- It requires no hyperparameter tuning
- It runs faster on large datasets
- It only works with continuous features
Correct answer: It captures non-linear relationships between variables
Mutual information measures any statistical dependency between variables, not just linear relationships.
Question 12: A model performs well on training data but poorly on test data. This is a sign of what?
- Data leakage from test to train
- Overfitting (Correct answer)
- Underfitting
- Class imbalance
Correct answer: Overfitting
Overfitting occurs when a model memorizes training noise and fails to generalize.
Question 13: Median imputation is often preferred over mean imputation when the column:
- Is skewed or has outliers (Correct answer)
- Has no missing values
- Is already scaled
- Is categorical
Correct answer: Is skewed or has outliers
The median is robust to skew and outliers, unlike the mean.
Question 14: What is the correct order of operations when preparing a dataset with missing values and features requiring scaling?
- Remove missing values and scaling simultaneously
- Apply feature scaling first, then impute missing values
- Scale only the complete cases and leave missing values as-is
- Impute missing values first, then apply feature scaling (Correct answer)
Correct answer: Impute missing values first, then apply feature scaling
Imputation must occur before scaling because most scaling methods cannot handle missing values and would either error or produce misleading results.
Question 15: What does a high recall but low precision indicate?
- Many false negatives
- No predictions made
- Many false positives but few false negatives (Correct answer)
- Perfect classification
Correct answer: Many false positives but few false negatives
High recall catches most positives, but low precision means many predicted positives are wrong.
Question 16: What does the 'naive' assumption in Naive Bayes refer to?
- The dataset is small
- The model ignores prior probabilities
- All classes are equally likely
- Features are conditionally independent given the class (Correct answer)
Correct answer: Features are conditionally independent given the class
Naive Bayes assumes features are conditionally independent given the class label.
Question 17: L1 (Lasso) regularization performs feature selection by:
- Shrinking some coefficients exactly to zero (Correct answer)
- Adding interaction terms
- Scaling features to unit variance
- Squaring all coefficients
Correct answer: Shrinking some coefficients exactly to zero
L1 penalty can drive coefficients to exactly zero, effectively removing those features.
Question 18: In 5-fold cross-validation, what fraction of data is used for testing in each fold?
- About 20% (Correct answer)
- About 80%
- About 5%
- About 50%
Correct answer: About 20%
With 5 folds, each test fold is 1/5 (20%) while 80% trains the model.
Question 19: A column of customer ages contains the value 999 for records where age was unknown. What is this an example of?
- A valid outlier
- A categorical variable
- A normalized feature
- A sentinel value encoding missing data (Correct answer)
Correct answer: A sentinel value encoding missing data
Sentinel values like 999 are placeholders that secretly encode missing or unknown data and should be converted to NaN before analysis.
Question 20: What is the main advantage of using the F-beta score with beta=2 instead of the standard F1 score?
- It equally weights precision and recall
- It eliminates the need for a classification threshold
- It places more emphasis on recall than precision (Correct answer)
- It places more emphasis on precision than recall
Correct answer: It places more emphasis on recall than precision
F-beta with beta=2 weighs recall twice as heavily as precision, making it ideal when missing positive cases is costlier than false alarms.
Question 21: Which measure describes how spread out data values are around the mean?
- Median
- Standard deviation (Correct answer)
- Percentile
- Mode
Correct answer: Standard deviation
Standard deviation quantifies the dispersion of data around the average.
Question 22: In simple exponential smoothing, the smoothing parameter α controls:
- The weight applied to the seasonal component only
- The rate at which the influence of older observations decays (Correct answer)
- The strength of the trend component in the forecast
- The total number of observations included in each forecast
Correct answer: The rate at which the influence of older observations decays
α (between 0 and 1) determines how quickly past observations lose influence — a higher α means more weight on recent data and faster adaptation to changes.
Question 23: What does the term 'data leakage' refer to in the context of data preparation?
- Data being lost during transfer between systems
- Information from outside the training set improperly influencing model building (Correct answer)
- Sensitive data being exposed to unauthorized users
- Memory overflow during large dataset processing
Correct answer: Information from outside the training set improperly influencing model building
Data leakage occurs when information that would not be available at prediction time is used during training, leading to overly optimistic performance estimates.
Question 24: Which measure is most appropriate for describing the spread of a skewed distribution?
- Interquartile range (Correct answer)
- Mode
- Standard deviation
- Mean
Correct answer: Interquartile range
The IQR is robust to skew and outliers, making it suitable for non-symmetric distributions.
Question 25: What does the 'I' stand for in the ARIMA model?
- Independent
- Integrated (Correct answer)
- Interaction
- Interval
Correct answer: Integrated
The 'I' in ARIMA stands for 'Integrated,' referring to the differencing of the series to achieve stationarity before fitting the AR and MA components.
Question 26: In a hypothesis test, a p-value of 0.03 with a significance level of 0.05 leads to which conclusion?
- Accept the alternative hypothesis as proven
- Fail to reject the null hypothesis
- Reject the null hypothesis (Correct answer)
- The test is inconclusive
Correct answer: Reject the null hypothesis
Since the p-value (0.03) is less than the significance level (0.05), we reject the null hypothesis.
Question 27: What is the primary purpose of regularization in supervised learning models such as Ridge and Lasso regression?
- To speed up the training process
- To increase the number of features in the model
- To prevent overfitting by penalizing large coefficients (Correct answer)
- To convert categorical variables into numerical ones
Correct answer: To prevent overfitting by penalizing large coefficients
Regularization adds a penalty term to the loss function that discourages overly complex models with large coefficient values.
Question 28: How does SARIMA differ from a standard ARIMA model?
- SARIMA does not require the time series to be stationary
- SARIMA adds seasonal AR, differencing, and MA terms to capture seasonal patterns (Correct answer)
- SARIMA uses machine learning instead of classical statistical methods
- SARIMA is exclusively designed for monthly or quarterly data only
Correct answer: SARIMA adds seasonal AR, differencing, and MA terms to capture seasonal patterns
SARIMA extends ARIMA by adding seasonal autoregressive, integrated, and moving average terms (SARIMA(p,d,q)(P,D,Q)m) to model periodic seasonal patterns alongside non-seasonal structure.
Question 29: A residual plot showing a clear curved pattern suggests:
- The model fits perfectly
- The data is normally distributed
- The linear model may be misspecified (Correct answer)
- There are no outliers
Correct answer: The linear model may be misspecified
Structure in residuals indicates the linear form fails to capture the true relationship.
Question 30: A data scientist applies the Bonferroni correction when conducting 20 simultaneous hypothesis tests at alpha = 0.05. What is the adjusted significance level for each individual test?
- 0.05
- 0.001
- 0.01
- 0.0025 (Correct answer)
Correct answer: 0.0025
The Bonferroni correction divides the overall significance level by the number of tests: 0.05 / 20 = 0.0025.
Question 31: A medical research team develops a model to screen for a rare but aggressive form of cancer. The consequences of failing to identify a person who has the cancer (a false negative) are far more severe than mistakenly flagging a healthy person for additional testing (a false positive). Which evaluation metric is most critical to maximize for this model?
- Precision
- Specificity
- Recall (Sensitivity) (Correct answer)
- Accuracy
Correct answer: Recall (Sensitivity)
Recall, also known as Sensitivity, measures the model's ability to correctly identify all actual positive cases (True Positives / (True Positives + False Negatives)). In this medical scenario, minimizing false negatives is the top priority to ensure patients with cancer receive timely treatment. Therefore, maximizing recall is the most critical objective.
Question 32: Which chart type is most appropriate for showing the part-to-whole composition of a single category at one point in time?
- Pie chart (Correct answer)
- Line chart
- Scatter plot
- Box plot
Correct answer: Pie chart
A pie chart shows how parts make up a whole for a single categorical breakdown.
Question 33: A correlation coefficient of -0.85 between two variables indicates:
- A weak negative relationship
- A strong positive relationship
- No relationship
- A strong negative linear relationship (Correct answer)
Correct answer: A strong negative linear relationship
Values near -1 indicate a strong inverse linear association.
Question 34: Mean imputation of missing values can distort which property of a feature?
- Its column name
- Its data type only
- Its variance, which becomes artificially reduced (Correct answer)
- The number of rows
Correct answer: Its variance, which becomes artificially reduced
Replacing missing values with the mean shrinks variance and weakens correlations.
Question 35: In a confusion matrix, what does the 'recall' metric measure?
- The proportion of actual positives correctly identified (Correct answer)
- The harmonic mean of precision and recall
- The proportion of predicted positives that are correct
- The overall accuracy of the model
Correct answer: The proportion of actual positives correctly identified
Recall measures the ability of a model to find all relevant positive instances out of the total actual positives.
Question 36: Which supervised learning algorithm constructs a series of if-then rules by recursively partitioning the feature space?
- DBSCAN
- Decision Tree (Correct answer)
- Principal Component Analysis
- K-Means Clustering
Correct answer: Decision Tree
Decision trees split data recursively using feature thresholds to create interpretable if-then decision rules.
Question 37: You apply a log transformation to a right-skewed feature containing zero values. What problem will this cause?
- It converts the feature to categorical
- It guarantees a normal distribution
- It removes all outliers automatically
- log(0) is undefined/negative infinity (Correct answer)
Correct answer: log(0) is undefined/negative infinity
log(0) is undefined, so a constant (e.g., log1p) must be added before transforming zeros.
Question 38: A data scientist observes that two variables have a Pearson correlation coefficient of -0.92. What does this indicate?
- A weak negative linear relationship
- No linear relationship
- A strong negative linear relationship (Correct answer)
- A strong positive linear relationship
Correct answer: A strong negative linear relationship
A Pearson correlation of -0.92 indicates a strong negative linear relationship, meaning as one variable increases, the other decreases proportionally.
Question 39: What problem does stratified sampling in cross-validation specifically address?
- Eliminating multicollinearity
- Preserving class distribution across folds (Correct answer)
- Reducing computation time
- Temporal data leakage
Correct answer: Preserving class distribution across folds
Stratified sampling ensures each fold maintains the same proportion of each class as the full dataset, preventing biased evaluation from uneven splits.
Question 40: What is the primary purpose of regularization (L1/L2) in supervised models?
- Increase training accuracy
- Speed up gradient computation
- Reduce overfitting by penalizing large weights (Correct answer)
- Balance class labels
Correct answer: Reduce overfitting by penalizing large weights
Regularization adds a penalty on coefficient magnitude to discourage overly complex models that overfit.
Question 41: Which statistical test is most appropriate for determining whether two independent groups have significantly different means, assuming normally distributed data and unknown but equal variances?
- Independent samples t-test (Student's t-test) (Correct answer)
- Paired t-test
- Chi-square test
- Mann-Whitney U test
Correct answer: Independent samples t-test (Student's t-test)
The independent samples t-test (Student's t-test) is designed to compare the means of two separate, unrelated groups when the data is approximately normally distributed and the variances are assumed to be equal. It produces a t-statistic and p-value to assess statistical significance.
Question 42: Which scaling method is most robust when a feature contains many extreme outliers?
- Min-max scaling
- Standard scaling
- Unit-vector scaling
- RobustScaler (uses median and IQR) (Correct answer)
Correct answer: RobustScaler (uses median and IQR)
RobustScaler centers on the median and scales by the interquartile range, making it resistant to outliers.
Question 43: Which measure of central tendency is most heavily influenced by extreme outlier values in a dataset?
- Mode
- Mean (Correct answer)
- Median
- Range
Correct answer: Mean
The mean (average) is calculated by summing all values and dividing by the count, so a single extreme outlier can pull the mean significantly in one direction, unlike the median or mode.
Question 44: When cleaning a customer database, a data analyst discovers that the 'State' column contains inconsistencies such as 'CA', 'Calif.', and 'California'. What is the most appropriate data cleaning step to address this issue?
- Standardize the categorical values to a single format. (Correct answer)
- Apply feature scaling to the 'State' column.
- Remove the 'State' column from the dataset.
- Impute missing values using the mode.
Correct answer: Standardize the categorical values to a single format.
The issue described is one of inconsistent formatting for a categorical feature. The correct approach is to standardize these values into a single, consistent format (e.g., converting all variations to 'CA'). This ensures that records are grouped correctly during analysis and that the feature is treated as a single category by machine learning models. Imputation is for missing data, removing the column would cause information loss, and feature scaling applies to numerical data.
Question 45: Why should scaling parameters be computed only on the training set, then applied to the test set?
- To prevent data leakage from the test set (Correct answer)
- To convert types
- To make the test set larger
- To remove duplicates
Correct answer: To prevent data leakage from the test set
Fitting the scaler on training data only prevents test-set statistics from leaking into the model and inflating performance.
Question 46: A data scientist obtains a 95% confidence interval of (12.3, 18.7) for the population mean. Which interpretation is correct?
- 95% of sample means fall in this interval
- The sample mean has a 95% chance of being between 12.3 and 18.7
- There is a 95% probability the population mean lies in this interval
- If we repeated sampling many times, 95% of constructed intervals would contain the true mean (Correct answer)
Correct answer: If we repeated sampling many times, 95% of constructed intervals would contain the true mean
A 95% confidence interval means that 95% of intervals constructed from repeated samples would capture the true population parameter.
Question 47: A key difference between PCA and feature selection is that PCA:
- Requires the target variable
- Always keeps the original interpretable features
- Creates new combined features rather than keeping original ones (Correct answer)
- Cannot reduce dimensions
Correct answer: Creates new combined features rather than keeping original ones
PCA produces transformed components, whereas selection retains a subset of original features.
Question 48: Which of the following is a supervised learning task?
- Reducing dimensions with t-SNE
- Predicting house prices from labeled sales data (Correct answer)
- Grouping customers into segments without labels
- Detecting communities in a graph
Correct answer: Predicting house prices from labeled sales data
Supervised learning uses labeled outputs, such as known house prices, to train a model.
Question 49: When communicating uncertainty in an estimate, which visual element is most appropriate?
- Error bars or confidence interval bands (Correct answer)
- A bold single point with no range
- A larger font
- A brighter color
Correct answer: Error bars or confidence interval bands
Error bars or confidence bands explicitly show the range of uncertainty around an estimate.
Question 50: Target leakage in feature engineering occurs when a feature:
- Has missing values
- Has too few categories
- Is perfectly scaled
- Contains information not available at prediction time (Correct answer)
Correct answer: Contains information not available at prediction time
Leakage happens when a feature reveals future or target information unavailable during real prediction.
Question 51: What is the purpose of applying a Box-Cox transformation during data preparation?
- To convert categorical variables into numerical format
- To reduce the number of features through dimensionality reduction
- To make a skewed distribution more closely approximate a normal distribution (Correct answer)
- To remove all outliers from the dataset automatically
Correct answer: To make a skewed distribution more closely approximate a normal distribution
Box-Cox transformation applies a power transformation that reduces skewness and helps data better satisfy normality assumptions required by many statistical methods.
Question 52: Stripping leading/trailing whitespace and fixing inconsistent capitalization in a text column is part of:
- Feature scaling
- Cross-validation
- Text/string cleaning (Correct answer)
- Dimensionality reduction
Correct answer: Text/string cleaning
Trimming whitespace and normalizing case are string-cleaning steps.
Question 53: In the bias-variance tradeoff, a model that is too simple to capture the underlying pattern exhibits:
- Low bias, low variance
- High variance, low bias
- High variance, high bias
- High bias, low variance (Correct answer)
Correct answer: High bias, low variance
An overly simple model underfits, showing high bias and low variance.
Question 54: A data scientist decides to convert a continuous 'Age' feature into a categorical 'Age_Group' feature (e.g., '18-25', '26-40', '41-60', '61+'). Which of the following is a primary benefit of this technique, known as binning?
- It is the only way to handle missing values in the original continuous feature.
- It increases the precision of the data by adding more information.
- It helps to capture non-linear relationships when using linear models. (Correct answer)
- It guarantees that the new feature will have a normal distribution.
Correct answer: It helps to capture non-linear relationships when using linear models.
Binning (or discretization) can help linear models capture non-linear relationships. [1] For example, the effect of age on a target variable might not be linear. By converting age into bins, a linear model can assign a different weight to each age group, effectively modeling a non-linear pattern without using a more complex model. [16]
Question 55: Which metric is most appropriate for evaluating a classification model when the dataset has a severe class imbalance (e.g., 95% negative, 5% positive)?
- R-squared
- Accuracy
- Mean Absolute Error
- Area Under the Precision-Recall Curve (AUPRC) (Correct answer)
Correct answer: Area Under the Precision-Recall Curve (AUPRC)
AUPRC focuses on the performance of the minority class and is more informative than accuracy when class distribution is highly imbalanced.
Question 56: What is the bias-variance tradeoff in supervised learning?
- The choice between supervised and unsupervised approaches
- The balance between the number of features and the number of samples
- The balance between a model's ability to fit training data closely and its ability to generalize to new data (Correct answer)
- The tradeoff between training speed and prediction accuracy
Correct answer: The balance between a model's ability to fit training data closely and its ability to generalize to new data
The bias-variance tradeoff describes how reducing bias (underfitting) often increases variance (overfitting) and vice versa.
Question 57: A precision-recall curve is generally preferred over an ROC curve when:
- The model is a regressor
- There are no false positives
- The positive class is rare (highly imbalanced) (Correct answer)
- Classes are perfectly balanced
Correct answer: The positive class is rare (highly imbalanced)
PR curves are more informative than ROC curves when the positive class is rare.
Question 58: When should you use a stacked area chart instead of multiple line charts?
- When comparing exact individual values precisely
- When showing a single series
- When plotting categorical data
- When emphasizing cumulative totals and part-to-whole over time (Correct answer)
Correct answer: When emphasizing cumulative totals and part-to-whole over time
Stacked area charts emphasize how components accumulate into a total across time.
Question 59: To avoid data leakage, feature scaling parameters should be fit on:
- The test set only
- The entire dataset before splitting
- Each fold's validation data
- The training set only, then applied to the test set (Correct answer)
Correct answer: The training set only, then applied to the test set
Fitting scalers only on training data prevents test information from leaking into the pipeline.
Question 60: Using the IQR method, an outlier is typically a value beyond which boundary?
- Above the median
- More than 1 standard deviation away
- Below Q1 - 1.5*IQR or above Q3 + 1.5*IQR (Correct answer)
- Below the mean
Correct answer: Below Q1 - 1.5*IQR or above Q3 + 1.5*IQR
The Tukey IQR rule flags points more than 1.5 times the interquartile range beyond the first or third quartile.
Question 61: Which Python library is primarily used for creating dataframes and performing data manipulation?
- Matplotlib
- Pandas (Correct answer)
- Scikit-learn
- TensorFlow
Correct answer: Pandas
Pandas provides the DataFrame structure and a rich set of functions for data cleaning, transformation, and analysis.
Data Science Certification Exam
The Data Science Certification Exam exam validates essential knowledge and skills required for certification or licensure in this field.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds