Spark Architecture and Cluster Management Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Spark Architecture and Cluster Management flashcards as text
What is the role of the Spark Driver in a Spark application?
Answer: Hosts the SparkContext, maintains the DAG, and coordinates task scheduling across executors
The Driver program runs the main() function, creates SparkContext, builds the DAG, and communicates with the cluster manager to schedule tasks.
What are Spark Executors?
Answer: JVM processes launched by the cluster manager that execute tasks and store RDD data
Executors are JVM processes running on worker nodes that execute tasks assigned by the Driver and cache RDD partitions in memory.
Which cluster manager is natively built into Apache Spark for standalone deployments?
Answer: Spark Standalone Mode
Spark's built-in standalone cluster manager provides a simple way to deploy Spark without YARN, Mesos, or Kubernetes.
What is a Spark Stage?
Answer: A phase of the DAG separated by wide transformations (shuffle boundaries)
A Stage is a set of tasks that can be computed without shuffling data; stage boundaries are determined by wide transformations.
What is a Spark Task?
Answer: A unit of work sent to one executor to process one data partition
A Task is the smallest unit of work in Spark, processing one data partition and running on a single executor core.
In Spark's DAG (Directed Acyclic Graph), what does each node represent?
Answer: An RDD or DataFrame with transformations applied
In Spark's DAG, each node represents an RDD or DataFrame, and edges represent transformations applied between them.