Mixed Deck — All Data Engineering Topics Flashcards
100 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 20 Mixed Deck — All Data Engineering Topics flashcards as text
Which US law protects the privacy of patient health information?
Answer: HIPAA
HIPAA governs protected health information in the US.
The 'timestamp' snapshot strategy in dbt relies on:
Answer: An updated-at column that reliably changes on every update
The timestamp strategy trusts an updated_at field to detect modified records.
A query spilling to disk during a sort or aggregation usually indicates:
Answer: Insufficient memory for the operation
When working memory is exceeded, the engine spills intermediate data to disk, slowing the query.
Which architecture uses both a batch layer and a speed layer to serve queries?
Answer: Lambda architecture
Lambda architecture combines a batch layer for accuracy with a speed layer for low latency.
Which key type uniquely identifies a row in a dimension table within the warehouse, independent of the source system?
Answer: Surrogate key
A surrogate key is a warehouse-generated integer that uniquely identifies dimension rows regardless of source identifiers.
An accumulating snapshot fact table is most appropriate for modeling:
Answer: A process with well-defined milestones, like order fulfillment stages
Accumulating snapshots track a process with multiple milestone dates, updating the same row as stages complete.
Which technique replaces sensitive values with non-sensitive surrogates that can be mapped back via a secure vault?
Answer: Tokenization
Tokenization substitutes data with tokens stored separately.
Which Parquet feature allows query engines to skip entire row groups based on filter conditions without reading the data?
Answer: Row group filtering using min/max statistics
Parquet stores min/max statistics for each row group and column chunk, allowing query engines to skip entire row groups when filter conditions cannot match any row in that group.
A pipeline reprocesses the entire source table every night instead of only changed rows. Which technique would reduce this load?
Answer: Incremental loading with change data capture
Incremental loading via CDC processes only inserted, updated, or deleted records since the last run.
What does 'denormalization' do to a database design?
Answer: Intentionally adds redundancy to improve read/query performance
Denormalization adds redundant data to reduce joins and speed up read-heavy queries.
Which of the following is a core mechanism for achieving fault tolerance in a distributed data processing system?
Answer: Data replication
Fault tolerance is the ability of a system to continue operating despite component failures. [21] A primary technique to achieve this is data replication, which involves storing multiple copies of data on different nodes or machines. [18] If one node holding a piece of data fails, the system can retrieve the data from another replica, ensuring availability and preventing data loss. [15]
A rapidly changing attribute (like customer age band) is best handled by splitting it into a:
Answer: Mini-dimension (junk or demographic dimension)
Volatile attributes are moved into a mini-dimension to avoid exploding the main dimension's row count.
What does data lineage describe in a pipeline context?
Answer: The path data takes from source through transformations to final outputs
Lineage traces how data flows and transforms across systems, aiding debugging and impact analysis.
What is the role of a message broker like Apache Kafka in a data architecture?
Answer: It decouples producers and consumers by buffering streams of events
Kafka acts as a durable, scalable buffer that decouples data producers from consumers via topic-based event streams.
What is the main reason to rotate access keys and credentials regularly?
Answer: Limit the window of exposure if a credential is compromised
Regular rotation shortens how long a leaked credential remains valid.
Which metric describes the time delay between data creation and its availability for use?
Answer: Data latency
Data latency measures the delay between when data is generated and when it becomes usable.
Which technique mitigates skew by adding random prefixes to hot keys?
Answer: Salting
Salting spreads a hot key across multiple reducers by appending random values.
What does 'speculative execution' do in a distributed batch framework?
Answer: Launches duplicate copies of slow tasks to finish faster
Speculative execution runs backup copies of straggler tasks and uses whichever finishes first.
Which Airflow concept controls how many task instances can run concurrently for a single DAG?
Answer: concurrency (max_active_tasks)
The concurrency (max_active_tasks) setting limits how many task instances run at once within a DAG.
What problem does a presigned URL solve in cloud object storage?
Answer: Granting temporary, scoped access to a private object without sharing credentials
A presigned URL grants time-limited access to a specific object without exposing account credentials.