Master of Data science Big Data 1 Flashcards
6 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Master of Data science Big Data 1 flashcards as text
Which file format is optimized for columnar storage and is commonly used in big data ecosystems for efficient query performance?
Answer: Parquet
Parquet is a columnar storage format designed for big data processing. It stores data by column rather than row, enabling efficient compression and query performance when only specific columns need to be read.
In the Lambda architecture for big data processing, which layer handles real-time data processing with low latency?
Answer: Speed layer
The speed layer (also called the streaming layer) in Lambda architecture processes incoming data in real time with low latency to provide up-to-date results, compensating for the high latency of the batch layer.
Which algorithm is commonly used for distributed machine learning to split gradient computation across multiple nodes?
Answer: Distributed Gradient Boosting
Distributed Gradient Boosting (as implemented in frameworks like XGBoost and LightGBM) splits the gradient computation and histogram building across multiple nodes, enabling scalable ensemble learning on large datasets.
What is the primary role of Apache ZooKeeper in a big data cluster?
Answer: Distributed coordination and configuration management
Apache ZooKeeper provides distributed coordination services such as configuration management, naming, synchronization, and group membership, allowing distributed applications like Hadoop and Kafka to maintain consistency across nodes.
Which concept in big data describes the phenomenon where storing and processing costs decrease as the volume of data increases per unit?
Answer: Economies of scale
Economies of scale in big data refers to the cost efficiency gained as data volume grows — distributed systems like Hadoop and cloud storage reduce the per-unit cost of storage and computation at higher volumes.
In Apache Spark, which operation triggers the actual execution of a transformation chain on an RDD or DataFrame?
Answer: collect()
Spark uses lazy evaluation, meaning transformations like map() and filter() are not executed immediately. An action such as collect() triggers the actual computation by submitting a job to the Spark cluster.