Test Flashcards
9 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 9 Test flashcards as text
Spark supports which cluster managers?
Answer: All of the above
Apache Spark is designed to be flexible regarding cluster management. It supports running on various cluster managers, including its own Standalone cluster manager, Apache Mesos, and Hadoop YARN, allowing users to choose the best environment for their needs. This broad compatibility makes Spark adaptable to different infrastructure setups.
Which of the following statements about Spark MLlib is correct?
Answer: It is the scalable machine learning library which delivers efficiencies
Spark MLlib is Apache Spark's scalable machine learning library, offering a wide range of machine learning algorithms and utilities. It is designed for efficiency and can process large datasets in a distributed manner, making it suitable for big data analytics. Its focus on scalability and efficiency is a core characteristic.
RDDs are immutable and fault-tolerant.
Answer: True
RDDs (Resilient Distributed Datasets) are indeed immutable, meaning once created, their contents cannot be changed. They are also fault-tolerant because they can be rebuilt from their lineage of transformations if a partition is lost, ensuring data integrity. These properties are fundamental to Spark's reliability and performance.
Which algorithm is not a solution for the regression problem?
Answer: Logistic Regression
Logistic Regression is primarily used for classification problems, where the goal is to predict a categorical outcome (e.g., yes/no, true/false). While it has 'regression' in its name, it models the probability of a binary outcome rather than predicting a continuous value, unlike the other options which are suitable for regression tasks. This distinction is crucial in machine learning.
Which of the following statements about Spark R is correct?
Answer: It allows data scientists to analyze large datasets and interactively run jobs
SparkR is an R package that provides a lightweight frontend to use Apache Spark from R. It enables R users to interact with Spark, perform data analysis on large datasets, and leverage Spark's distributed computing capabilities within their familiar R environment. This integration empowers data scientists to scale their R workflows.
Which of the following statements regarding DataFrame is correct?
Answer: DataFrames provide a more user-friendly API than RDDs.
DataFrames offer a higher-level, more structured, and user-friendly API compared to RDDs, resembling tables in a relational database. They provide schema information and allow for optimized execution plans, making data manipulation and querying more intuitive and efficient for many use cases. This improved usability is a key advantage over raw RDDs.
Which of the following statements about Spark Shell is correct?
Answer: All of the above
The Spark Shell is an interactive environment (REPL) that allows users to quickly experiment with Spark code, test applications, and perform interactive data analysis. It supports reading from various data sources and facilitates running Spark applications directly from the command line. Its versatility makes it an essential tool for development and exploration.
Is MLlib a deprecated library?
Answer: No
MLlib is not deprecated; it is actively maintained and continues to be Apache Spark's primary machine learning library. While there's also ML (Spark ML), which is a newer, DataFrame-based API, MLlib (the RDD-based API) still exists and is used, especially for certain algorithms or custom implementations. Both libraries coexist and serve different use cases.
On RDD, the read operation is
Answer: Either fine-grained or coarse-grained
RDDs support both coarse-grained and fine-grained read operations, though they are primarily optimized for coarse-grained transformations. Coarse-grained operations apply to the entire dataset (e.g., map, filter), while fine-grained operations allow access to specific elements or partitions. This flexibility allows for various data access patterns.