Big Data Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Big Data flashcards as text
Which technique does Apache Spark use to optimize query execution plans in Spark SQL?
Answer: Catalyst optimizer
Spark SQL uses the Catalyst optimizer to apply rule-based and cost-based optimizations on logical and physical query plans.
In stream processing, what does 'exactly-once' semantics guarantee?
Answer: Each message produces one and only one output effect
Exactly-once semantics ensures that despite failures and retries, each input message affects system state only once.
Which data format is columnar and widely used in the Hadoop ecosystem for efficient analytical queries?
Answer: Parquet
Parquet is a columnar storage format that enables efficient column pruning and predicate pushdown for analytical queries.
What is the primary role of Apache ZooKeeper in a distributed big data system?
Answer: Distributed coordination and consensus
ZooKeeper provides distributed coordination services such as leader election, configuration management, and distributed locks.
Which windowing strategy in stream processing assigns each event to all windows that contain its timestamp?
Answer: Sliding window
Sliding windows overlap, so a single event can belong to multiple windows depending on the slide interval.
In MapReduce, the combiner function is best described as a:
Answer: Local reducer that runs on mapper output before the shuffle
The combiner acts as a mini-reducer on the mapper side, aggregating local output to reduce data transferred during shuffle.
Which consistency model does Apache Cassandra use by default for reads and writes?
Answer: Eventual consistency
Cassandra defaults to eventual consistency, prioritizing availability and partition tolerance over immediate consistency.