Optimizing Query Performance Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Optimizing Query Performance flashcards as text
Data skew in a distributed join causes which symptom?
Answer: A few tasks run far longer than the rest
Skew concentrates rows on a few keys, overloading the tasks that process them.
Which technique mitigates skew on a hot join key?
Answer: Salting the key to spread rows across partitions
Salting appends a random component to distribute a hot key over more partitions.
A covering index improves a query because it:
Answer: Contains all columns the query needs, avoiding table lookups
When the index includes every referenced column, the engine answers from the index alone.
SELECT * in an analytical query on columnar storage is discouraged mainly because it:
Answer: Forces reading all columns, increasing I/O
Selecting every column negates columnar's ability to skip unused columns.
Materializing an expensive, frequently-run aggregation into a summary table is best described as:
Answer: A materialized view / precomputation tradeoff
Precomputing results trades storage and refresh cost for faster reads.
Which condition typically prevents an index from being used on a column?
Answer: Wrapping the column in a function in the WHERE clause
Functions applied to a column make the predicate non-sargable, blocking index use.
When two large tables are joined, which strategy avoids broadcasting and instead shuffles both sides by key?
Answer: Sort-merge / shuffle hash join
For two large inputs, both are repartitioned by the join key and merged or hashed.