Big Data Technologies Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Big Data Technologies flashcards as text
Which Apache Spark optimization technique avoids shuffling by co-partitioning two datasets on the same key?
Answer: Sort-merge join with bucketing
Bucketing pre-partitions data on a join key so that matching partitions are co-located, eliminating the shuffle step during joins.
In Google's original MapReduce paper, what is the purpose of the 'combiner' function?
Answer: To perform partial aggregation on the mapper's output before it is sent to the reducer, reducing network traffic
The combiner acts as a mini-reducer on each mapper node, aggregating intermediate key-value pairs locally to minimize data shuffled over the network.
What is Apache Iceberg's primary advantage over a traditional Hive partitioned table?
Answer: Iceberg provides ACID transactions, hidden partitioning, and time-travel on large analytic tables without full table scans on metadata
Iceberg adds a metadata layer that enables ACID guarantees, partition evolution, and time-travel queries without modifying the underlying Parquet/ORC files.
Which consistency model does Apache Cassandra use when the consistency level is set to 'QUORUM'?
Answer: A majority (more than half) of replicas in the replication factor must acknowledge
QUORUM requires acknowledgment from a majority of replicas, balancing consistency and availability.
In stream processing, what does 'exactly-once semantics' guarantee?
Answer: Each record is processed exactly once, with no duplicates and no data loss
Exactly-once semantics ensures every record affects the output exactly one time, combining idempotent writes with transactional commits or distributed snapshots.
What problem does Apache Kafka's 'log compaction' feature solve?
Answer: It retains only the most recent value for each key, enabling changelog-style topics to be replayed to current state
Log compaction ensures that for each unique key, the latest value is always retained, making the topic suitable for state reconstruction.
Which of the following best describes the concept of 'schema-on-read' used in data lakes?
Answer: Data is stored raw without enforced schema, and structure is applied at query time
Schema-on-read allows raw data to be ingested without upfront transformation, deferring schema enforcement to when the data is queried.