Introduction To Data Engineering Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Introduction To Data Engineering flashcards as text
In ELT, where does the transformation step occur compared to ETL?
Answer: Inside the destination system after loading raw data
ELT loads raw data into the target first, then transforms it using the destination's compute power.
What is data partitioning primarily used for?
Answer: Improving query performance and manageability by splitting data into segments
Partitioning divides large datasets into segments so queries can scan less data and run faster.
Which file format is columnar and commonly used in big data analytics?
Answer: Parquet
Parquet is a columnar storage format optimized for analytical queries on large datasets.
What does 'idempotency' mean for a data pipeline operation?
Answer: Running it multiple times produces the same result as running it once
An idempotent operation yields the same outcome no matter how many times it is executed.
Which tool is widely used to orchestrate and schedule data workflows?
Answer: Apache Airflow
Apache Airflow orchestrates workflows as directed acyclic graphs (DAGs) with scheduling and monitoring.
What is a primary key in a relational table?
Answer: A column or set of columns that uniquely identifies each row
A primary key uniquely identifies every row and cannot contain null or duplicate values.
What is the main benefit of data compression in storage systems?
Answer: Reduced storage costs and faster I/O for large datasets
Compression shrinks data size, lowering storage costs and often speeding up read/write operations.