Big Data Technologies Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Big Data Technologies flashcards as text
Which component of the Hadoop ecosystem is responsible for distributed storage?
Answer: HDFS
HDFS (Hadoop Distributed File System) is the storage layer that splits files into blocks and distributes them across cluster nodes.
In Apache Kafka, what is a 'consumer group' used for?
Answer: Allowing multiple consumers to read different partitions in parallel
A consumer group lets multiple consumers divide the partitions of a topic among themselves, enabling parallel and scalable consumption.
What does the CAP theorem state about distributed systems?
Answer: You can guarantee at most two of Consistency, Availability, and Partition tolerance
The CAP theorem states that a distributed system can guarantee only two of the three properties: Consistency, Availability, and Partition tolerance.
Which Apache project provides a SQL-like interface for querying data stored in HDFS?
Answer: Hive
Apache Hive provides HiveQL, a SQL-like language that translates queries into MapReduce or Tez jobs over HDFS data.
In Spark, what is the primary advantage of using DataFrames over RDDs?
Answer: DataFrames enable Catalyst optimizer and Tungsten execution engine optimizations
DataFrames expose schema information that allows Spark's Catalyst optimizer and Tungsten execution engine to apply significant performance optimizations.
What is 'data skew' in the context of distributed big data processing?
Answer: Uneven distribution of data across partitions causing some tasks to run much longer
Data skew occurs when some partitions hold significantly more data than others, creating bottleneck tasks that slow down the entire job.
Which storage format is columnar and commonly used in the Hadoop ecosystem for analytical workloads?
Answer: Parquet
Parquet is a columnar storage format that provides efficient compression and encoding, making it well-suited for analytical queries on large datasets.