Data Transformation & Cleaning Flashcards
9 cards from real ADE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 9 Data Transformation & Cleaning flashcards as text
What is data transformation in pipelines?
Answer: Converting data format
Data transformation is a key stage in data pipelines where raw data is converted, cleaned, aggregated, and enriched to meet the requirements of the target system or analytical use case. This can involve changing data types, joining datasets, calculating new fields, or standardizing values. The goal is to make the data suitable for accurate and meaningful analysis.
Which technique removes duplicate records?
Answer: Deduplication
Deduplication is a data quality technique used to identify and remove redundant copies of data within a dataset. Duplicate records can skew analytical results, waste storage space, and lead to inconsistencies. By ensuring that each unique entity or record is represented only once, deduplication improves data accuracy, integrity, and efficiency for downstream processes.
What is data normalization?
Answer: Restructuring data to reduce redundancy
Data normalization is a database design technique used to organize tables in a relational database to minimize data redundancy and improve data integrity. It involves breaking down large tables into smaller, related tables and defining relationships between them. This structured approach ensures that data is stored logically, efficiently, and consistently, reducing update anomalies.
Which method fills missing data values?
Answer: Imputation
Imputation is a data preprocessing technique used to fill in missing data values with substituted values. Various methods exist, such as mean, median, mode imputation, or more advanced statistical models, to estimate and replace the missing entries. This ensures that datasets are complete and can be used for analysis without discarding valuable records, which might otherwise lead to biased results.
What is the goal of data cleansing?
Answer: To correct errors
Data cleansing, also known as data scrubbing, is the process of detecting and correcting or removing corrupt, inaccurate, or irrelevant records from a dataset. Its primary goal is to improve data quality by ensuring data is accurate, consistent, and complete. This is vital for reliable analysis, accurate reporting, and sound decision-making, as flawed data can lead to incorrect insights.
Which tool is commonly used for data transformation?
Answer: Apache Spark
Apache Spark is a powerful open-source distributed processing engine designed for big data workloads. It provides robust APIs for data transformation, including filtering, aggregation, joining, and cleaning, making it a go-to tool for ETL (Extract, Transform, Load) processes in data engineering. Its in-memory processing capabilities enable fast and efficient data manipulation across large datasets.
Why is schema validation important during transformation?
Answer: To ensure data format correctness
Schema validation during data transformation is crucial for maintaining data quality and integrity. It verifies that incoming data conforms to the expected structure, data types, and constraints of the target schema. This prevents malformed or inconsistent data from entering downstream systems, which could lead to errors, incorrect analyses, or system failures.
What is data enrichment?
Answer: Adding relevant external data
Data enrichment is the process of enhancing existing data by integrating it with additional, relevant information from external sources. This process adds more context and value to the original dataset, making it more comprehensive and useful for analysis or decision-making. For example, adding demographic data to customer records or geographic information to location data.
What is the role of data profiling?
Answer: To analyze data quality and structure
Data profiling is the process of examining the data available in an existing information source and collecting statistics and information about that data. Its role is to assess data quality, identify patterns, discover relationships, and understand the structure and content of the data. This helps in identifying anomalies, inconsistencies, and potential issues before transformation or loading.