โ† All Data Engineering Flashcard Decks

Distributed Data Processing Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Distributed Data Processing flashcards as text
  1. What is the benefit of partition pruning in a distributed query engine?

    Answer: It skips reading partitions irrelevant to the filter

    Partition pruning avoids scanning data that cannot match the query predicate.

  2. In the lambda architecture, what is the speed layer responsible for?

    Answer: Low-latency views of recent data

    The speed layer serves real-time results while the batch layer recomputes accurate views.

  3. Which factor most directly limits the scalability of a global ORDER BY in distributed SQL?

    Answer: All data must funnel through a single sort/merge step

    Total ordering forces a global merge that becomes a bottleneck at scale.

  4. What does eventual consistency guarantee?

    Answer: Replicas converge to the same value if updates stop

    Given no new writes, all replicas eventually reflect the same data.

  5. Why are columnar formats often paired with predicate pushdown?

    Answer: The engine can skip blocks using column statistics

    Min/max statistics per column block let the engine skip non-matching data.

  6. In a consistent hashing ring, what happens when one node is removed?

    Answer: Only its keys remap to the next node

    Consistent hashing limits remapping to the departing node's keys, minimizing disruption.

  7. What is the main reason to colocate compute with data in a cluster?

    Answer: Reduce network transfer by processing data locally

    Data locality minimizes network I/O by running tasks where the data already resides.