Distributed Data Processing Flashcards
7 cards from real Data Engineering practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Distributed Data Processing flashcards as text
In a MapReduce job, what is the primary purpose of the shuffle phase?
Answer: To group and transfer mapper output to reducers by key
The shuffle phase sorts mapper output and routes records with the same key to the same reducer.
What problem does data skew most directly cause in a distributed join?
Answer: A few overloaded tasks slow the whole stage
Skew concentrates many records on one key, making a few tasks far slower than the rest.
Which Spark operation triggers actual computation rather than just building the DAG?
Answer: count()
count() is an action, while filter, map, and select are lazy transformations.
A broadcast join is most appropriate when:
Answer: One table is small enough to fit in each executor's memory
Broadcasting a small table to every node avoids shuffling the large table.
What does HDFS block replication primarily provide?
Answer: Fault tolerance against node failure
Multiple replicas let the cluster recover data if a DataNode fails.
In Spark, what is a 'narrow' transformation?
Answer: Each output partition depends on one input partition
Narrow transformations like map require no data movement across partitions.
Why is checkpointing useful in long-running streaming jobs?
Answer: It persists state so the job can recover after failure
Checkpoints save progress and state, enabling recovery without reprocessing everything.