← All Data Engineering Flashcard Decks

Optimizing Query Performance Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Optimizing Query Performance flashcards as text
  1. Too many small files ("small files problem") hurts query performance mainly by:

    Answer: Adding per-file open/metadata overhead

    Each tiny file incurs listing and open costs that dominate when files are numerous.

  2. Compaction in a data lake table addresses performance by:

    Answer: Merging many small files into fewer larger ones

    Compaction consolidates small files to reduce scan overhead.

  3. Caching a hot dimension table in memory across a cluster helps most when:

    Answer: It is reused across many queries and fits in memory

    Caching pays off for small, frequently reused datasets that fit in memory.

  4. A query spilling to disk during a sort or aggregation usually indicates:

    Answer: Insufficient memory for the operation

    When working memory is exceeded, the engine spills intermediate data to disk, slowing the query.

  5. Choosing a partition column with very high cardinality (e.g., user_id) often backfires because it:

    Answer: Creates a huge number of tiny partitions

    Over-partitioning produces excessive small partitions and metadata overhead.

  6. LIMIT with an ORDER BY on a non-indexed column still requires the engine to:

    Answer: Sort the full result set before returning rows

    Without a supporting index, the full set must be sorted to find the top rows.

  7. Bucketing a table by the join key primarily improves performance by:

    Answer: Avoiding a shuffle since matching keys are co-located

    Pre-bucketing on the join key lets engines join corresponding buckets without reshuffling.