CompTIA Data+ (DA0-001) — Questions and Answers
Question 1: What is role-based access control (RBAC) in data systems?
- A security model granting data access based on a user's organizational role (Correct answer)
- A database replication strategy
- An encryption method for sensitive data
- A method for assigning random access permissions
Correct answer: A security model granting data access based on a user's organizational role
RBAC assigns permissions to roles rather than individuals, granting users access to data based on their job function within the organization.
Question 2: What is the purpose of a choropleth map?
- Using color intensity to represent data values across geographic regions (Correct answer)
- Showing flight paths between cities
- Animating data changes over time
- Displaying 3D topographic data
Correct answer: Using color intensity to represent data values across geographic regions
A choropleth map shades geographic regions with colors proportional to a data variable, enabling geographic comparisons.
Question 3: What is Snowflake's multi-cluster, shared data architecture?
- A peer-to-peer distributed computing model
- Multiple databases sharing one compute cluster
- Separate storage and compute layers allowing multiple independent query engines (Correct answer)
- A single cluster serving all workloads
Correct answer: Separate storage and compute layers allowing multiple independent query engines
Snowflake separates storage from compute, allowing multiple virtual warehouses to independently access the same data without contention.
Question 4: What is a Type I error in hypothesis testing?
- Using the wrong statistical test
- Accepting an incorrect alternative hypothesis
- Rejecting a true null hypothesis (Correct answer)
- Failing to reject a false null hypothesis
Correct answer: Rejecting a true null hypothesis
A Type I error (false positive) occurs when a true null hypothesis is incorrectly rejected.
Question 5: What is a KPI in the context of data visualization?
- Key Process Integration
- Kernel Processing Index
- Key Performance Indicator (Correct answer)
- Known Pattern Insight
Correct answer: Key Performance Indicator
A KPI (Key Performance Indicator) is a measurable value that demonstrates how effectively a company is achieving key business objectives.
Question 6: What is the purpose of a data retention policy?
- To define how long data should be kept before it is archived or deleted (Correct answer)
- To maximize data storage for analytics
- To prevent data from being modified after creation
- To specify who has access to sensitive data
Correct answer: To define how long data should be kept before it is archived or deleted
A data retention policy specifies how long different categories of data are kept, balancing business needs, legal requirements, and storage costs.
Question 7: What is a star schema in data warehousing?
- A schema using only a single table
- A schema with multiple fact tables connected to each other
- A database schema designed for transactional processing
- A central fact table connected to multiple dimension tables (Correct answer)
Correct answer: A central fact table connected to multiple dimension tables
A star schema has a central fact table containing measurable metrics connected to multiple denormalized dimension tables, optimized for analytical queries.
Question 8: Which aggregate function returns the total number of rows in a result set?
- COUNT() (Correct answer)
- MAX()
- AVG()
- SUM()
Correct answer: COUNT()
COUNT() returns the number of rows that match the specified condition, or total rows when used with *.
Question 9: What is a snowflake schema in data warehousing?
- A variation of the star schema where dimension tables are normalized into sub-dimensions (Correct answer)
- A schema with no fact tables
- A star schema with only one dimension
- A schema designed for cold storage data
Correct answer: A variation of the star schema where dimension tables are normalized into sub-dimensions
A snowflake schema normalizes dimension tables into multiple related tables, reducing redundancy at the cost of more complex joins.
Question 10: Which type of machine learning is used in recommendation systems where an agent learns by receiving rewards or penalties?
- Supervised learning
- Reinforcement learning (Correct answer)
- Semi-supervised learning
- Unsupervised learning
Correct answer: Reinforcement learning
Reinforcement learning trains an agent to make decisions by rewarding desired actions and penalizing undesired ones, making it well-suited for sequential decision-making tasks.
Question 11: What is the California Consumer Privacy Act (CCPA)?
- An FTC regulation on digital advertising
- A federal law requiring encryption of all consumer data
- A national cybersecurity standard for data protection
- A California state law giving consumers rights to know, delete, and opt out of the sale of their personal data (Correct answer)
Correct answer: A California state law giving consumers rights to know, delete, and opt out of the sale of their personal data
CCPA is a California privacy law granting consumers rights to know what personal data is collected, request deletion, and opt out of data sales.
Question 12: Which type of machine learning involves training a model on labeled data to predict outcomes for new, unseen data?
- Self-supervised learning
- Reinforcement learning
- Unsupervised learning
- Supervised learning (Correct answer)
Correct answer: Supervised learning
Supervised learning uses labeled training data where the correct output is known, allowing the model to learn mappings from inputs to outputs.
Question 13: What does GDPR stand for and how does it affect US companies?
- General Data Protection Regulation — applies to US companies handling EU residents' data (Correct answer)
- Government Data Protection Registry — applies to federal agencies only
- Global Data Privacy Requirements — applies to all multinational companies
- General Data Privacy Rule — applies only in Europe
Correct answer: General Data Protection Regulation — applies to US companies handling EU residents' data
GDPR is the EU's data protection regulation that applies to any organization, including US companies, that processes personal data of EU residents.
Question 14: What is the purpose of a CDN (Content Delivery Network) in analytics applications?
- To process big data queries
- To store analytical models
- To encrypt data at rest
- To reduce latency by caching content closer to users geographically (Correct answer)
Correct answer: To reduce latency by caching content closer to users geographically
A CDN distributes cached content across geographically dispersed servers to reduce latency and improve load times for end users.
Question 15: What is a slowly changing dimension (SCD) in data warehousing?
- A dimension whose values change very slowly due to database performance issues
- A real-time dimension updated every second
- A dimension table with few records
- A technique for managing historical changes to dimension attributes over time (Correct answer)
Correct answer: A technique for managing historical changes to dimension attributes over time
SCD (Slowly Changing Dimension) is a method for handling changes to dimension attributes over time, preserving history of how values changed.
Question 16: What is OLAP in business intelligence?
- Operational Lookup and Aggregation Protocol
- Online Analytical Processing for multidimensional analysis (Correct answer)
- Online Linear Analytics Processing
- Object-Level Analytics Platform
Correct answer: Online Analytical Processing for multidimensional analysis
OLAP (Online Analytical Processing) enables fast, multidimensional analysis of large datasets by organizing data into cubes for slicing, dicing, and drilling.
Question 17: Which algorithm builds an ensemble of decision trees using random subsets of features and training samples?
- Random Forest (Correct answer)
- Naive Bayes
- K-Nearest Neighbors
- Gradient Boosting
Correct answer: Random Forest
Random Forest creates multiple decision trees using random feature subsets and bootstrap samples, then aggregates their predictions to improve accuracy and reduce overfitting.
Question 18: What does R-squared (R²) measure in regression analysis?
- The slope of the regression line
- The statistical significance of the model
- The proportion of variance in the dependent variable explained by the model (Correct answer)
- The number of predictors in the model
Correct answer: The proportion of variance in the dependent variable explained by the model
R-squared measures the proportion of variance in the dependent variable that is explained by the independent variable(s) in the model.
Question 19: What is Bayes' Theorem used for?
- Calculating the median of a dataset
- Updating probability estimates based on new evidence (Correct answer)
- Measuring data spread around the mean
- Testing differences between two populations
Correct answer: Updating probability estimates based on new evidence
Bayes' Theorem calculates the conditional probability of an event by updating prior beliefs with new evidence.
Question 20: What is prescriptive analytics?
- Analyzing past performance data
- Predicting what will happen next quarter
- Describing the current state of business
- Recommending specific actions to achieve desired outcomes based on data models (Correct answer)
Correct answer: Recommending specific actions to achieve desired outcomes based on data models
Prescriptive analytics uses optimization and simulation algorithms to recommend specific actions organizations should take to achieve desired outcomes.
Question 21: Which Python library is most commonly used for statistical data visualization?
- NumPy
- Pandas
- Seaborn (Correct answer)
- Scikit-learn
Correct answer: Seaborn
Seaborn is built on Matplotlib and provides a high-level interface for drawing attractive and informative statistical graphics.
Question 22: What does a PRIMARY KEY constraint enforce in a relational database?
- Referential integrity
- Unique, non-null values per row (Correct answer)
- Column data types
- Default column values
Correct answer: Unique, non-null values per row
A PRIMARY KEY ensures every row has a unique, non-null identifier in that column or set of columns.
Question 23: What is a master data management (MDM) system?
- A system for managing database administrator accounts
- A technology for creating a single, authoritative source of truth for critical business data (Correct answer)
- An ETL tool for moving data between systems
- A database backup and recovery system
Correct answer: A technology for creating a single, authoritative source of truth for critical business data
MDM creates and maintains a single, consistent, and accurate record of critical business entities like customers, products, and suppliers across the organization.
Question 24: What is data deduplication?
- Creating redundant database copies for failover
- Encrypting duplicate copies of data
- The process of identifying and removing duplicate records from datasets (Correct answer)
- Backing up data in multiple locations
Correct answer: The process of identifying and removing duplicate records from datasets
Data deduplication identifies and eliminates duplicate records in datasets to ensure each entity appears only once, improving data quality.
Question 25: Which SQL window function assigns a unique rank to each row within a partition?
- LAG()
- PARTITION BY
- ROW_NUMBER() (Correct answer)
- SUM() OVER
Correct answer: ROW_NUMBER()
ROW_NUMBER() assigns a unique sequential integer to each row within a partition, with no gaps or ties.
Question 26: What is the purpose of a validation set in machine learning?
- To train the model's parameters
- To tune hyperparameters and select the best model during development (Correct answer)
- To augment the training data
- To provide a final unbiased evaluation of the model
Correct answer: To tune hyperparameters and select the best model during development
A validation set is used during model development to tune hyperparameters and compare different models, separate from the test set used for final evaluation.
Question 27: Which type of chart is best for showing data trends over time?
- Bar chart
- Pie chart
- Line chart (Correct answer)
- Scatter plot
Correct answer: Line chart
A line chart connects data points chronologically, making it ideal for visualizing trends, patterns, and changes over time.
Question 28: Which AWS service is a fully managed data warehouse solution?
- Amazon Redshift (Correct answer)
- Amazon RDS
- Amazon Aurora
- Amazon DynamoDB
Correct answer: Amazon Redshift
Amazon Redshift is AWS's fully managed, petabyte-scale cloud data warehouse optimized for analytics workloads.
Question 29: What is a confidence interval?
- The probability that a hypothesis is correct
- The margin of error in data collection
- A range of values likely to contain the true population parameter (Correct answer)
- The range within which the sample mean falls
Correct answer: A range of values likely to contain the true population parameter
A confidence interval provides a range of plausible values for a population parameter with a specified level of confidence (e.g., 95%).
Question 30: In Power BI, what is a 'slicer' used for?
- Filtering dashboard visuals interactively by selected values (Correct answer)
- Slicing columnar data in a table
- Cutting data into training and test sets
- Dividing reports into pages
Correct answer: Filtering dashboard visuals interactively by selected values
A slicer in Power BI is an interactive filter control that lets users segment all visuals on a report page by selected values.
Question 31: What are the five dimensions commonly used to measure data quality?
- Format, Frequency, Freshness, Fidelity, Flexibility
- Speed, Scale, Security, Scope, Structure
- Accuracy, Completeness, Consistency, Timeliness, Uniqueness (Correct answer)
- Availability, Reliability, Validity, Visibility, Volume
Correct answer: Accuracy, Completeness, Consistency, Timeliness, Uniqueness
The five core data quality dimensions are Accuracy (correct values), Completeness (no missing data), Consistency (no contradictions), Timeliness (up to date), and Uniqueness (no duplicates).
Question 32: Which color scale is recommended for sequential data with a meaningful zero point?
- Rainbow color scale
- Qualitative palette
- Diverging color scale
- Sequential single-hue scale (Correct answer)
Correct answer: Sequential single-hue scale
A sequential single-hue scale uses varying lightness of one color to represent ordered data, making magnitude differences clear.
Question 33: Which distribution is characterized by a long right tail and is often seen in income data?
- Uniform distribution
- Right-skewed (positive skew) distribution (Correct answer)
- Bimodal distribution
- Normal distribution
Correct answer: Right-skewed (positive skew) distribution
A right-skewed distribution has a tail extending to higher values, common in income data where most people earn moderate amounts but a few earn very high incomes.
Question 34: What is a data lake?
- A cloud-based ETL pipeline
- A relational database optimized for analytics
- A data warehouse with real-time capabilities
- A centralized repository storing raw data in native format at any scale (Correct answer)
Correct answer: A centralized repository storing raw data in native format at any scale
A data lake stores structured, semi-structured, and unstructured data in its raw format, enabling flexible analysis without predefined schemas.
Question 35: What is Apache Hadoop primarily used for?
- Relational database management
- Distributed storage and batch processing of large datasets (Correct answer)
- In-memory data analytics
- Real-time stream processing
Correct answer: Distributed storage and batch processing of large datasets
Apache Hadoop is a distributed computing framework that uses HDFS for storage and MapReduce for batch processing of large datasets across clusters.
Question 36: Which of the following is an example of an unsupervised learning algorithm?
- Random forest
- K-means clustering (Correct answer)
- Support vector machine
- Linear regression
Correct answer: K-means clustering
K-means clustering is unsupervised because it groups data points based on similarity without using labeled training examples.
Question 37: What is 'chart junk' in data visualization?
- Charts with too many data points
- Visual elements that add clutter without conveying information (Correct answer)
- Broken or corrupted chart files
- Incorrectly formatted axis labels
Correct answer: Visual elements that add clutter without conveying information
Chart junk refers to unnecessary visual elements like decorative patterns, 3D effects, and gridlines that clutter charts without adding meaning.
Question 38: What is a Gantt chart primarily used for in data projects?
- Project scheduling and timeline management (Correct answer)
- Displaying statistical correlations
- Mapping data lineage
- Visualizing database schemas
Correct answer: Project scheduling and timeline management
A Gantt chart displays project tasks as horizontal bars along a timeline, showing start dates, durations, and dependencies.
Question 39: What does a scatter plot primarily show?
- Trends over time
- Frequency of categories
- Relationship between two numeric variables (Correct answer)
- Proportions of a whole
Correct answer: Relationship between two numeric variables
A scatter plot displays the correlation or relationship between two continuous numeric variables using dots.
Question 40: Which SQL command permanently removes a table and its data from the database?
- TRUNCATE
- REMOVE
- DELETE
- DROP (Correct answer)
Correct answer: DROP
DROP TABLE removes the entire table definition and all its data permanently from the database.
Question 41: What is the purpose of a data catalog in modern BI?
- A searchable inventory of data assets with metadata, definitions, and lineage information (Correct answer)
- A catalog of BI report templates
- A system for archiving old reports
- A physical storage system for data backups
Correct answer: A searchable inventory of data assets with metadata, definitions, and lineage information
A data catalog helps users discover, understand, and trust data assets by providing metadata, business definitions, data lineage, and usage information.
Question 42: What does the bias-variance tradeoff describe in machine learning?
- The tradeoff between model accuracy and training speed
- The tradeoff between the number of features and sample size
- The balance between precision and recall
- The balance between underfitting (high bias) and overfitting (high variance) (Correct answer)
Correct answer: The balance between underfitting (high bias) and overfitting (high variance)
The bias-variance tradeoff describes how increasing model complexity reduces bias (underfitting) but increases variance (overfitting), requiring a balance for optimal generalization.
Question 43: What is ad hoc reporting in business intelligence?
- Scheduled reports generated automatically at fixed intervals
- Reports generated by AI without user input
- Custom, on-demand reports created by users to answer specific questions (Correct answer)
- Standard reports built by IT departments
Correct answer: Custom, on-demand reports created by users to answer specific questions
Ad hoc reporting allows business users to create custom, one-time reports on demand without relying on IT, using self-service BI tools.
Question 44: In machine learning, what does the term 'feature engineering' refer to?
- Transforming raw data into meaningful input variables for a model (Correct answer)
- Selecting the best algorithm for a task
- Evaluating model performance on a test set
- Deploying a trained model to production
Correct answer: Transforming raw data into meaningful input variables for a model
Feature engineering is the process of using domain knowledge to create, transform, or select variables (features) from raw data to improve model performance.
Question 45: What chart type is most appropriate for comparing parts of a whole across multiple categories?
- Box plot
- Stacked bar chart (Correct answer)
- Scatter plot
- Waterfall chart
Correct answer: Stacked bar chart
A stacked bar chart shows how each category's total is composed of subcategories, enabling part-to-whole comparisons.
Question 46: What is predictive analytics?
- Describing current business conditions
- Analyzing what happened in the past
- Using historical data and statistical models to forecast future outcomes (Correct answer)
- Creating real-time dashboards
Correct answer: Using historical data and statistical models to forecast future outcomes
Predictive analytics uses historical data, statistical algorithms, and machine learning to identify the likelihood of future outcomes.
Question 47: What does 'data sovereignty' mean?
- The authority of the data governance team over data policies
- The concept that data is subject to the laws of the country where it is stored or collected (Correct answer)
- A company's ownership of all its data assets
- An individual's right to control their personal data
Correct answer: The concept that data is subject to the laws of the country where it is stored or collected
Data sovereignty means data is governed by the laws and regulations of the country in which it is physically located or collected.
Question 48: What does a NULL value represent in a database?
- An unknown or missing value (Correct answer)
- Zero
- An empty string
- A negative number
Correct answer: An unknown or missing value
NULL represents the absence of a value or an unknown value in a database field.
Question 49: Which SQL clause is used to filter rows after aggregation?
- GROUP BY
- WHERE
- ORDER BY
- HAVING (Correct answer)
Correct answer: HAVING
HAVING filters groups created by GROUP BY, whereas WHERE filters individual rows before aggregation.
Question 50: What does a box plot (box-and-whisker plot) display?
- Correlation between two variables
- Frequency of each category
- Distribution summary including median, quartiles, and outliers (Correct answer)
- Mean and standard deviation only
Correct answer: Distribution summary including median, quartiles, and outliers
A box plot shows the five-number summary (minimum, Q1, median, Q3, maximum) and visually highlights outliers.
Question 51: What SQL statement is used to retrieve data from a database?
- SELECT (Correct answer)
- INSERT
- DELETE
- UPDATE
Correct answer: SELECT
SELECT is the SQL DML statement used to query and retrieve data from one or more database tables.
Question 52: What is the interquartile range (IQR)?
- The average of Q1 and Q3
- The standard deviation of the middle 50% of data
- The difference between max and min values
- The difference between the 75th and 25th percentiles (Correct answer)
Correct answer: The difference between the 75th and 25th percentiles
The IQR is Q3 minus Q1, representing the spread of the middle 50% of data and used to identify outliers.
Question 53: What is overfitting in a machine learning model?
- The model learns training data too well and fails to generalize to new data (Correct answer)
- The model requires too much computational power
- The model is too simple to capture patterns in the data
- The model performs poorly on both training and test data
Correct answer: The model learns training data too well and fails to generalize to new data
Overfitting occurs when a model memorizes training data noise and details, resulting in high training accuracy but poor performance on unseen data.
Question 54: In k-fold cross-validation, what happens to the data?
- The data is bootstrapped k times to create new training sets
- The data is split once into training and test sets
- The model is trained k times on the full dataset
- The data is divided into k equal parts; each fold serves as the test set once while the rest train the model (Correct answer)
Correct answer: The data is divided into k equal parts; each fold serves as the test set once while the rest train the model
K-fold cross-validation partitions data into k folds, cycling through each fold as the test set while training on the remaining k-1 folds, then averaging the results.
Question 55: What does 'drill down' mean in OLAP analysis?
- Filtering data by a specific dimension
- Moving from a summary view to a more detailed level of data (Correct answer)
- Combining multiple data sources
- Removing data from a report
Correct answer: Moving from a summary view to a more detailed level of data
Drilling down navigates from a higher-level summary (e.g., annual sales) to a more granular level (e.g., monthly or daily sales).
Question 56: What is PII in the context of data governance?
- Primary Integration Interface
- Publicly Integrated Information
- Personally Identifiable Information that can identify a specific individual (Correct answer)
- Personal Identification Index
Correct answer: Personally Identifiable Information that can identify a specific individual
PII (Personally Identifiable Information) is any data that can be used to identify a specific individual, such as name, SSN, email, or biometric data.
Question 57: What does standard deviation measure?
- The range between minimum and maximum values
- The most frequent value in a dataset
- The average distance of data points from the mean (Correct answer)
- The middle value in a dataset
Correct answer: The average distance of data points from the mean
Standard deviation quantifies the dispersion of data points around the mean, measuring how spread out values are.
Question 58: What does the SQL DISTINCT keyword do?
- Sorts results alphabetically
- Filters NULL values
- Removes duplicate rows from the result set (Correct answer)
- Limits the number of rows returned
Correct answer: Removes duplicate rows from the result set
DISTINCT eliminates duplicate rows from the query result, returning only unique value combinations.
Question 59: Which principle of data visualization does Edward Tufte's 'data-ink ratio' emphasize?
- Maximizing the proportion of ink used to display data (Correct answer)
- Using bright colors for emphasis
- Adding gridlines for readability
- Including as many chart types as possible
Correct answer: Maximizing the proportion of ink used to display data
Tufte's data-ink ratio advocates removing all non-essential chart elements to maximize the proportion of ink conveying actual data.
Question 60: What is the purpose of cross-validation in machine learning?
- To increase model training speed
- To visualize model performance
- To assess how well a model generalizes to independent datasets (Correct answer)
- To remove outliers from training data
Correct answer: To assess how well a model generalizes to independent datasets
Cross-validation evaluates model generalization by training and testing on different subsets of data, reducing overfitting risk.
Question 61: What is the role of a data steward in a BI organization?
- To manage cloud infrastructure
- To build ETL pipelines
- To manage data quality, definitions, and governance policies for assigned data domains (Correct answer)
- To write SQL queries for analysts
Correct answer: To manage data quality, definitions, and governance policies for assigned data domains
A data steward oversees data quality, maintains business definitions, and enforces data governance policies for specific data domains within the organization.
Question 62: What is a streaming analytics pipeline?
- A dashboard that refreshes daily
- A batch job that runs every hour
- A type of ETL process for historical data
- Real-time processing of continuous data streams as events arrive (Correct answer)
Correct answer: Real-time processing of continuous data streams as events arrive
A streaming analytics pipeline processes data continuously as it arrives, enabling real-time insights rather than waiting for batch processing windows.
Question 63: What does the p-value represent in hypothesis testing?
- The effect size of the tested variable
- The probability of observing results as extreme as the data if the null hypothesis is true (Correct answer)
- The confidence interval width
- The probability the null hypothesis is true
Correct answer: The probability of observing results as extreme as the data if the null hypothesis is true
The p-value is the probability of obtaining results at least as extreme as observed, assuming the null hypothesis is true.
Question 64: What is the primary advantage of using ensemble methods like bagging?
- They train faster than single models on large datasets
- They reduce the number of hyperparameters to tune
- They eliminate the need for feature engineering
- They combine multiple models to reduce variance and improve predictive performance (Correct answer)
Correct answer: They combine multiple models to reduce variance and improve predictive performance
Bagging (Bootstrap Aggregating) trains multiple models on different bootstrap samples and averages their predictions, reducing variance and improving overall model stability.
Question 65: What is a fact table in a data warehouse?
- A table containing only textual descriptions
- A table containing database configuration settings
- A central table storing measurable, numeric business metrics and foreign keys to dimensions (Correct answer)
- A lookup table for product codes
Correct answer: A central table storing measurable, numeric business metrics and foreign keys to dimensions
A fact table stores quantitative business metrics (like sales revenue or order quantity) along with foreign keys linking to dimension tables.
Question 66: What does 'data mesh' refer to in modern data architecture?
- A network topology for data centers
- A centralized data lake with mesh indexing
- A decentralized approach where domain teams own and manage their own data as products (Correct answer)
- A type of mesh network for IoT sensors
Correct answer: A decentralized approach where domain teams own and manage their own data as products
Data mesh is a decentralized architectural approach that treats data as a product, with domain-oriented teams owning their own data pipelines.
Question 67: What is the main purpose of a dashboard in data analytics?
- To store raw data efficiently
- To automate data collection
- To provide a real-time overview of key metrics in one view (Correct answer)
- To replace detailed analytical reports
Correct answer: To provide a real-time overview of key metrics in one view
A dashboard consolidates key performance indicators and metrics into a single, at-a-glance interface for monitoring business health.
Question 68: What does a chi-square test assess?
- The association between two categorical variables (Correct answer)
- The normality of a dataset
- The difference between two means
- The correlation between two continuous variables
Correct answer: The association between two categorical variables
A chi-square test of independence evaluates whether there is a statistically significant association between two categorical variables.
Question 69: In a neural network, what is the role of an activation function?
- To normalize input data before training
- To reduce the number of parameters in the model
- To introduce non-linearity so the network can learn complex patterns (Correct answer)
- To initialize the weights of the network
Correct answer: To introduce non-linearity so the network can learn complex patterns
Activation functions introduce non-linearity into neural networks, enabling them to learn and represent complex, non-linear relationships in data.
Question 70: Which chart type is best for showing the distribution of a single continuous variable?
- Line chart
- Bar chart
- Histogram (Correct answer)
- Pie chart
Correct answer: Histogram
A histogram displays the frequency distribution of continuous data by grouping values into bins.
Question 71: What does a z-score of 2.0 indicate?
- The value is twice the median
- The value is 2 standard deviations above the mean (Correct answer)
- The value has a 2% probability of occurring
- The value is 2 units below the mean
Correct answer: The value is 2 standard deviations above the mean
A z-score of 2.0 means the data point is 2 standard deviations above the mean of the distribution.
Question 72: What is a treemap visualization used for?
- Mapping geographic data
- Displaying hierarchical data as nested rectangles sized by value (Correct answer)
- Showing decision tree algorithms
- Tracking project timelines
Correct answer: Displaying hierarchical data as nested rectangles sized by value
A treemap displays hierarchical data as nested rectangles, where rectangle size represents a quantitative variable.
Question 73: What is the mean of the dataset: 4, 8, 6, 10, 12?
- 10
- 8 (Correct answer)
- 7
- 9
Correct answer: 8
The mean is calculated as (4+8+6+10+12)/5 = 40/5 = 8.
Question 74: Which technique is used to prevent overfitting by adding a penalty term to the loss function?
- Bagging
- Regularization (Correct answer)
- Boosting
- Normalization
Correct answer: Regularization
Regularization adds a penalty (L1 or L2) to the loss function to discourage the model from fitting noise and reduce overfitting.
Question 75: What is the 'right to be forgotten' in data governance?
- A database feature for purging old transaction records
- A legal right for individuals to request deletion of their personal data from an organization's systems (Correct answer)
- An employee's right to delete their work history
- A policy for automatically deleting old analytics reports
Correct answer: A legal right for individuals to request deletion of their personal data from an organization's systems
The right to be forgotten (right to erasure) allows individuals to request that organizations delete their personal data, recognized under GDPR and similar regulations.
Question 76: What is the Central Limit Theorem (CLT)?
- The sampling distribution of the mean approaches normal as sample size increases (Correct answer)
- Larger datasets always have smaller standard deviations
- All populations follow a normal distribution
- The mean always equals the median in a dataset
Correct answer: The sampling distribution of the mean approaches normal as sample size increases
The CLT states that the sampling distribution of the sample mean approximates a normal distribution as sample size increases, regardless of population distribution.
Question 77: What is a heat map used to represent?
- Magnitude of values across two dimensions using color intensity (Correct answer)
- Geographic temperature data only
- Network connections between nodes
- Categorical comparisons over time
Correct answer: Magnitude of values across two dimensions using color intensity
A heat map encodes data values as color intensity across a matrix, making it easy to spot patterns and outliers.
Question 78: What does 'data at rest' refer to in security?
- Data temporarily cached in memory
- Data stored in databases, file systems, or storage media that is not actively moving (Correct answer)
- Data waiting to be processed in a queue
- Data that has not been accessed recently
Correct answer: Data stored in databases, file systems, or storage media that is not actively moving
Data at rest refers to inactive data stored persistently in databases, file systems, or storage media, as opposed to data in transit or in use.
Question 79: What is data tokenization?
- Converting data into blockchain tokens
- Splitting datasets into training tokens for ML
- Replacing sensitive data with non-sensitive placeholder tokens while preserving format (Correct answer)
- Encoding data in base64 format
Correct answer: Replacing sensitive data with non-sensitive placeholder tokens while preserving format
Tokenization substitutes sensitive data (like credit card numbers) with non-sensitive tokens that have no exploitable value if intercepted.
Question 80: What is the difference between OLTP and OLAP systems?
- OLTP handles reporting; OLAP handles transactions
- OLTP is cloud-based; OLAP is on-premise
- OLTP handles real-time transactions; OLAP handles analytical queries on historical data (Correct answer)
- OLTP uses NoSQL; OLAP uses relational databases only
Correct answer: OLTP handles real-time transactions; OLAP handles analytical queries on historical data
OLTP (Online Transaction Processing) systems handle high-volume day-to-day transactions, while OLAP systems are optimized for complex analytical queries on historical data.
Question 81: What does a waterfall chart primarily illustrate?
- Statistical distributions
- Hierarchical data structures
- Cumulative effect of sequential positive and negative values (Correct answer)
- Time series trends
Correct answer: Cumulative effect of sequential positive and negative values
A waterfall chart shows how an initial value is affected by a series of positive and negative incremental changes to reach a final value.
Question 82: What is a foreign key in a relational database?
- A column that references the primary key of another table (Correct answer)
- A key that stores NULL values
- An index on multiple columns
- A key used for encryption
Correct answer: A column that references the primary key of another table
A foreign key establishes a referential link between a column in one table and the primary key of another table.
Question 83: What is a balanced scorecard in business analytics?
- A dashboard showing only financial KPIs
- A strategic management tool measuring performance across financial, customer, process, and learning dimensions (Correct answer)
- A SQL query performance benchmark
- A type of database normalization
Correct answer: A strategic management tool measuring performance across financial, customer, process, and learning dimensions
The balanced scorecard measures organizational performance across four perspectives: financial, customer, internal processes, and learning and growth.
Question 84: What is transfer learning in the context of deep learning?
- Converting a model from one framework to another
- Transferring data between databases for training
- Sending a model from one server to another for inference
- Reusing a pre-trained model on a new but related task (Correct answer)
Correct answer: Reusing a pre-trained model on a new but related task
Transfer learning leverages knowledge from a model trained on a large dataset and fine-tunes it for a new, related task, reducing training time and data requirements.
Question 85: What is a data quality rule?
- A regulatory requirement from a government agency
- A SQL constraint on a database column
- A company policy about data ownership
- A condition or constraint that data must satisfy to be considered fit for use (Correct answer)
Correct answer: A condition or constraint that data must satisfy to be considered fit for use
A data quality rule defines the criteria that data must meet to be considered accurate, complete, and valid for its intended business use.
Question 86: What is data masking used for?
- Replacing sensitive data with realistic but fictitious data to protect privacy (Correct answer)
- Compressing data for efficient storage
- Encrypting data during transmission
- Hiding charts from unauthorized users
Correct answer: Replacing sensitive data with realistic but fictitious data to protect privacy
Data masking replaces sensitive real data with structurally similar but fictional values, enabling development and testing without exposing actual sensitive data.
Question 87: What is a data mart?
- A type of OLTP database
- A subset of a data warehouse focused on a specific business function or department (Correct answer)
- A real-time streaming database
- A marketplace for buying and selling data
Correct answer: A subset of a data warehouse focused on a specific business function or department
A data mart is a focused subset of a data warehouse tailored to meet the specific reporting and analysis needs of a business department like marketing or finance.
Question 88: What type of correlation does a Pearson coefficient of -0.85 indicate?
- Weak positive correlation
- Strong positive correlation
- Strong negative correlation (Correct answer)
- No correlation
Correct answer: Strong negative correlation
A Pearson coefficient of -0.85 indicates a strong negative linear relationship where one variable increases as the other decreases.
Question 89: What is Delta Lake?
- An open-source storage layer that brings ACID transactions to data lakes (Correct answer)
- A columnar file format for Hadoop
- A real-time streaming platform
- A cloud storage service by AWS
Correct answer: An open-source storage layer that brings ACID transactions to data lakes
Delta Lake is an open-source storage layer that adds ACID transaction support, schema enforcement, and time travel capabilities to data lakes.
Question 90: What is a data dictionary?
- A database index for faster lookups
- A programming library for data manipulation
- A glossary of data science terms for new employees
- A centralized repository documenting data elements, definitions, formats, and relationships (Correct answer)
Correct answer: A centralized repository documenting data elements, definitions, formats, and relationships
A data dictionary documents metadata about data elements including field names, data types, definitions, acceptable values, and relationships.
Question 91: Which visualization tool is most commonly used for business intelligence dashboards in the US?
- Tableau (Correct answer)
- R ggplot2
- Matplotlib
- Excel pivot tables
Correct answer: Tableau
Tableau is the leading BI dashboard tool in the US enterprise market, known for its drag-and-drop interface and interactive dashboards.
Question 92: Which of the following best describes hyperparameter tuning?
- Increasing the size of the training data
- Removing irrelevant features from the dataset
- Adjusting model weights during backpropagation
- Searching for the best configuration settings that control the learning process (Correct answer)
Correct answer: Searching for the best configuration settings that control the learning process
Hyperparameter tuning involves searching for optimal settings (like learning rate, tree depth, or regularization strength) that are set before training and control how the model learns.
CompTIA Data+ (DA0-001)
CompTIA Data+ validates skills in data mining, analysis, visualization, and governance for early-career data and business intelligence professionals. It covers the full data lifecycle from data concepts and mining through analysis, visualization, and governance quality controls.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds