Optimizing Query Performance Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Optimizing Query Performance flashcards as text
Too many small files ("small files problem") hurts query performance mainly by:
Answer: Adding per-file open/metadata overhead
Each tiny file incurs listing and open costs that dominate when files are numerous.
Compaction in a data lake table addresses performance by:
Answer: Merging many small files into fewer larger ones
Compaction consolidates small files to reduce scan overhead.
Caching a hot dimension table in memory across a cluster helps most when:
Answer: It is reused across many queries and fits in memory
Caching pays off for small, frequently reused datasets that fit in memory.
A query spilling to disk during a sort or aggregation usually indicates:
Answer: Insufficient memory for the operation
When working memory is exceeded, the engine spills intermediate data to disk, slowing the query.
Choosing a partition column with very high cardinality (e.g., user_id) often backfires because it:
Answer: Creates a huge number of tiny partitions
Over-partitioning produces excessive small partitions and metadata overhead.
LIMIT with an ORDER BY on a non-indexed column still requires the engine to:
Answer: Sort the full result set before returning rows
Without a supporting index, the full set must be sorted to find the top rows.
Bucketing a table by the join key primarily improves performance by:
Answer: Avoiding a shuffle since matching keys are co-located
Pre-bucketing on the join key lets engines join corresponding buckets without reshuffling.