Big Data Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Big Data flashcards as text
Which strategy reduces shuffle data volume in Spark by replacing groupByKey with a more efficient alternative?
Answer: Replacing groupByKey() with reduceByKey()
reduceByKey() applies a combining function locally on each partition before shuffling, dramatically reducing the data transferred across the network.
In data lakehouse architecture, what role does a metadata layer like Apache Iceberg or Delta Lake serve?
Answer: It adds ACID transactions, schema evolution, and time travel to data lake files
Table formats like Iceberg and Delta Lake maintain transactional metadata over raw files, enabling ACID guarantees, schema evolution, and historical snapshots.
Which metric best quantifies the throughput of a big data streaming system?
Answer: Events processed per second
Events (or messages) processed per second directly measures how much data the system handles over time, representing throughput.
Apache Beam's unified programming model is designed to allow pipelines to run on multiple execution engines. What is the abstraction called that wraps these engines?
Answer: Runner
A Beam Runner (e.g., Dataflow Runner, Flink Runner) translates the portable Beam pipeline into the target engine's native execution model.
In a distributed join, which optimization avoids a shuffle by sending a small table to all worker nodes?
Answer: Broadcast join
A broadcast join replicates the smaller table to every executor so the larger table can be joined locally without shuffling either dataset.
Which characteristic of big data refers to the uncertainty and unreliability of data sources?
Answer: Veracity
Veracity addresses data quality concerns such as noise, bias, and trustworthiness inherent in large heterogeneous data sources.
A data engineer needs to process a 10 TB dataset nightly with a 4-hour SLA. Which processing paradigm is most appropriate?
Answer: Scheduled batch processing
A nightly batch job with a 4-hour window fits scheduled batch processing, which optimizes throughput over large datasets without requiring streaming infrastructure.