Big Data Technologies Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Big Data Technologies flashcards as text
What does 'exactly-once semantics' mean in a distributed streaming system?
Answer: Each message is processed and its effect is reflected exactly once, even in the presence of failures
Exactly-once semantics guarantee that despite failures and retries, each message's effect on state or output appears exactly once, requiring idempotency and transactional support.
Which of the following is a key advantage of using Apache Iceberg over plain Parquet files on a data lake?
Answer: Iceberg provides ACID transactions, hidden partitioning, and schema evolution without full rewrites
Apache Iceberg adds a transactional metadata layer over Parquet (or ORC) files, enabling ACID operations, time travel, and seamless schema/partition evolution.
In HDFS, what is the NameNode's primary function?
Answer: Maintaining the filesystem metadata and directory tree, tracking where blocks are stored
The NameNode is the master server that manages the HDFS namespace, tracking which DataNodes hold which blocks, but it does not store data itself.
What is the purpose of Apache Oozie in the Hadoop ecosystem?
Answer: Workflow scheduling and coordination of Hadoop jobs
Apache Oozie is a workflow scheduler that chains and coordinates Hadoop jobs (MapReduce, Hive, Pig, Spark) based on time or data availability triggers.
Which metric is most important when evaluating the performance of a big data storage system under write-heavy workloads?
Answer: Write throughput measured in MB/s or records/second
Write throughput (records or bytes per second) directly measures how quickly the system can ingest data, which is the primary bottleneck in write-heavy workloads.
What does 'partition pruning' mean in the context of big data query optimization?
Answer: Skipping the scan of partitions that cannot contain data matching the query's filter conditions
Partition pruning allows the query engine to read only the partitions relevant to the filter predicate, dramatically reducing I/O on large partitioned datasets.
In a Kappa Architecture, what replaces the batch layer found in Lambda Architecture?
Answer: Reprocessing historical data through the same stream processing pipeline with a replay mechanism
Kappa Architecture eliminates the separate batch layer by reprocessing historical data by replaying it through the unified stream processing pipeline, simplifying operations.