Distributed Data Processing Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Distributed Data Processing flashcards as text
What does 'speculative execution' do in a distributed batch framework?
Answer: Launches duplicate copies of slow tasks to finish faster
Speculative execution runs backup copies of straggler tasks and uses whichever finishes first.
In Spark, persisting an RDD/DataFrame with caching mainly helps when:
Answer: It is reused across multiple actions
Caching avoids recomputation when the same dataset feeds several actions.
What is the consequence of having too many small files on HDFS?
Answer: NameNode memory pressure from excessive metadata
Each file consumes NameNode metadata, so many tiny files strain the NameNode.
Which technique mitigates skew by adding random prefixes to hot keys?
Answer: Salting
Salting spreads a hot key across multiple reducers by appending random values.
In a stream processor, what is backpressure?
Answer: A slowdown signal when downstream cannot keep up
Backpressure throttles upstream producers when consumers fall behind, preventing overload.
Why does idempotency matter in at-least-once processing?
Answer: It makes duplicate processing produce the same result
Idempotent operations let duplicate deliveries occur without corrupting results.
What does a YARN ResourceManager primarily do?
Answer: Allocates cluster resources to applications
The ResourceManager schedules and allocates CPU/memory containers across the cluster.