โ† All Data Engineering Flashcard Decks

Distributed Data Processing Flashcards

7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Distributed Data Processing flashcards as text
  1. What does 'speculative execution' do in a distributed batch framework?

    Answer: Launches duplicate copies of slow tasks to finish faster

    Speculative execution runs backup copies of straggler tasks and uses whichever finishes first.

  2. In Spark, persisting an RDD/DataFrame with caching mainly helps when:

    Answer: It is reused across multiple actions

    Caching avoids recomputation when the same dataset feeds several actions.

  3. What is the consequence of having too many small files on HDFS?

    Answer: NameNode memory pressure from excessive metadata

    Each file consumes NameNode metadata, so many tiny files strain the NameNode.

  4. Which technique mitigates skew by adding random prefixes to hot keys?

    Answer: Salting

    Salting spreads a hot key across multiple reducers by appending random values.

  5. In a stream processor, what is backpressure?

    Answer: A slowdown signal when downstream cannot keep up

    Backpressure throttles upstream producers when consumers fall behind, preventing overload.

  6. Why does idempotency matter in at-least-once processing?

    Answer: It makes duplicate processing produce the same result

    Idempotent operations let duplicate deliveries occur without corrupting results.

  7. What does a YARN ResourceManager primarily do?

    Answer: Allocates cluster resources to applications

    The ResourceManager schedules and allocates CPU/memory containers across the cluster.