Databricks Certified Associate Developer for Apache Spark — Questions and Answers
Question 1: Which feature transformer in Spark ML converts text documents to TF-IDF feature vectors?
- Word2Vec
- HashingTF + IDF (Correct answer)
- NGram
- BertEncoder
Correct answer: HashingTF + IDF
HashingTF converts text tokens to term frequency vectors, and IDF scales by inverse document frequency to produce TF-IDF features.
Question 2: What is the recommended way to handle small file problems in Spark?
- Increase the number of partitions to create more smaller files
- Disable speculative execution to reduce file output
- Set spark.files.maxPartitionBytes to a very large value
- Use coalesce() or repartition() to merge small partitions before writing (Correct answer)
Correct answer: Use coalesce() or repartition() to merge small partitions before writing
Using coalesce() to reduce partitions before writing merges small files into larger ones, improving subsequent read performance.
Question 3: What is the difference between a Transformer and an Estimator in Spark ML?
- A Transformer is for feature engineering; an Estimator is for evaluation
- A Transformer applies a transformation to a DataFrame; an Estimator fits data to produce a Transformer (Correct answer)
- A Transformer trains on data; an Estimator applies a transformation
- Both are identical; the names are interchangeable
Correct answer: A Transformer applies a transformation to a DataFrame; an Estimator fits data to produce a Transformer
An Estimator has a fit() method that trains on data and returns a Transformer, which has a transform() method.
Question 4: Which method forces Spark to recompute an RDD and remove its cached data?
- rdd.evict()
- rdd.release()
- rdd.clear()
- rdd.unpersist() (Correct answer)
Correct answer: rdd.unpersist()
unpersist() removes the RDD from the cache, freeing the memory for other computations.
Question 5: How do you register a DataFrame as a temporary SQL view in Spark?
- df.toTable("name")
- df.createOrReplaceTempView("name") (Correct answer)
- df.registerTable("name")
- spark.registerView(df, "name")
Correct answer: df.createOrReplaceTempView("name")
createOrReplaceTempView() registers a DataFrame as a temporary view scoped to the current SparkSession.
Question 6: Which data type is used as a vertex identifier (VertexId) in GraphX?
- Long (64-bit integer) (Correct answer)
- Integer (32-bit)
- String
- UUID
Correct answer: Long (64-bit integer)
GraphX defines VertexId as a type alias for Long, providing a large namespace of unique IDs suitable for billions of vertices.
Question 7: Which method waits for a streaming query to finish in Spark?
- query.block()
- query.wait()
- query.join()
- query.awaitTermination() (Correct answer)
Correct answer: query.awaitTermination()
awaitTermination() blocks the driver program until the streaming query is stopped or encounters an error.
Question 8: Which DataFrame function is used to select specific columns?
- select() (Correct answer)
- filter()
- where()
- pick()
Correct answer: select()
select() returns a new DataFrame with only the specified columns.
Question 9: How does Spark achieve fault tolerance with RDDs?
- By replicating data across multiple nodes
- By writing all intermediate data to disk
- By maintaining checkpoints after every transformation
- By recomputing lost partitions using lineage information (Correct answer)
Correct answer: By recomputing lost partitions using lineage information
Spark achieves fault tolerance by recomputing lost RDD partitions using the recorded lineage of transformations.
Question 10: Which Spark SQL function is used to perform an inner join between two DataFrames?
- df1.merge(df2, on='key')
- df1.join(df2, 'key') (Correct answer)
- df1.combine(df2, 'key')
- df1.link(df2, 'key')
Correct answer: df1.join(df2, 'key')
df1.join(df2, 'key') performs a join between two DataFrames; the default join type is inner.
Question 11: What is the purpose of the BlockManager in Apache Spark?
- Blocks network access between executors for security
- Manages storage of data blocks (RDD partitions, shuffle data, broadcast variables) in memory and on disk (Correct answer)
- Manages HDFS block replication for Spark's data
- Manages disk I/O operations for reading and writing data files
Correct answer: Manages storage of data blocks (RDD partitions, shuffle data, broadcast variables) in memory and on disk
BlockManager is Spark's distributed storage system that manages how data (RDD partitions, shuffle blocks, broadcasts) is stored across memory and disk.
Question 12: What is a Spark application's relationship to Spark jobs?
- Applications and jobs are synonymous in Spark
- A job contains one or more applications
- An application contains one or more jobs, each triggered by an action (Correct answer)
- An application contains exactly one job
Correct answer: An application contains one or more jobs, each triggered by an action
A Spark application can trigger multiple jobs; each action (collect, count, save) creates one job consisting of stages and tasks.
Question 13: What are Spark Executors?
- Worker nodes in the cluster that manage data replication
- Processes on the driver node that compile the DAG into tasks
- Memory managers that handle executor JVM garbage collection
- JVM processes launched by the cluster manager that execute tasks and store RDD data (Correct answer)
Correct answer: JVM processes launched by the cluster manager that execute tasks and store RDD data
Executors are JVM processes running on worker nodes that execute tasks assigned by the Driver and cache RDD partitions in memory.
Question 14: What is the purpose of the `subgraph` operation in GraphX?
- To create a coarser hierarchical summary of the graph
- To extract a subset of vertices and edges that satisfy a given predicate (Correct answer)
- To merge two separate graphs into one larger graph
- To split a graph into two equal halves for parallel processing
Correct answer: To extract a subset of vertices and edges that satisfy a given predicate
The `subgraph` operation filters vertices and edges using user-supplied predicates, returning a new graph that contains only the elements satisfying both conditions.
Question 15: What is Catalyst in Apache Spark?
- A machine learning library
- A streaming engine
- The query optimizer for Spark SQL (Correct answer)
- A storage manager for DataFrames
Correct answer: The query optimizer for Spark SQL
Catalyst is Spark SQL's extensible query optimizer that transforms logical plans into optimized physical plans.
Question 16: What parameters are used to define window operation?
- Window length, sliding interval (Correct answer)
- State size, sliding interval
- None of the above
- State size, window length
Correct answer: Window length, sliding interval
Window operations in Spark Streaming allow computations over a sliding window of data. These operations are defined by two key parameters: the 'window length,' which determines the duration of the data included in the window, and the 'sliding interval,' which specifies how often the window slides. These parameters control how data is aggregated and processed over time.
Question 17: What is an accumulator in Apache Spark?
- A type of broadcast variable that accumulates data over time
- A shared variable that can be read and written by all executors
- A monitoring tool that tracks resource consumption across executors
- A write-only distributed counter or aggregator that executors can add to and the driver can read (Correct answer)
Correct answer: A write-only distributed counter or aggregator that executors can add to and the driver can read
Accumulators are write-only from executor perspective — tasks can add to them, but only the driver can read the accumulated value.
Question 18: What does the `triplets` property of a GraphX Graph return?
- An array of three graphs representing different structural views
- An RDD of EdgeTriplet objects containing source vertex, edge, and destination vertex attributes (Correct answer)
- Three separate RDDs for source, edge, and destination
- A tuple of (VertexRDD, EdgeRDD, PartitionRDD)
Correct answer: An RDD of EdgeTriplet objects containing source vertex, edge, and destination vertex attributes
The `triplets` property returns an RDD[EdgeTriplet[VD, ED]], where each triplet combines the source vertex, the connecting edge, and the destination vertex with all their attributes.
Question 19: What is a Spark Stage?
- A phase of the DAG separated by wide transformations (shuffle boundaries) (Correct answer)
- A group of tasks that can be executed in parallel without a shuffle
- A container for multiple Spark jobs
- A single task running on one partition
Correct answer: A phase of the DAG separated by wide transformations (shuffle boundaries)
A Stage is a set of tasks that can be computed without shuffling data; stage boundaries are determined by wide transformations.
Question 20: What does the printSchema() method do in Spark DataFrames?
- Prints all rows of the DataFrame
- Displays the first 20 rows
- Prints the schema of the DataFrame in a tree format (Correct answer)
- Exports the schema to a file
Correct answer: Prints the schema of the DataFrame in a tree format
printSchema() prints the schema of the DataFrame as a tree showing column names, types, and nullability.
Question 21: Which method is used to persist an RDD in Spark?
- store()
- memorize()
- save()
- cache() or persist() (Correct answer)
Correct answer: cache() or persist()
cache() and persist() are used to persist an RDD; cache() is a shorthand for persist(StorageLevel.MEMORY_ONLY).
Question 22: What is the purpose of the VectorAssembler in Spark MLlib?
- Reduces the dimensionality of features
- Scales feature values to a standard range
- Encodes categorical features as binary vectors
- Combines multiple feature columns into a single vector column (Correct answer)
Correct answer: Combines multiple feature columns into a single vector column
VectorAssembler merges multiple numeric columns into a single feature vector required by Spark ML algorithms.
Question 23: What distributed collection type does GraphX use internally to store vertex attributes?
- VertexRDD (Correct answer)
- HashSet
- DataFrame
- DenseVector
Correct answer: VertexRDD
GraphX uses VertexRDD[VD], which extends RDD[(VertexId, VD)], to store and index vertex attributes as key-value pairs.
Question 24: Which of the following statements about Spark Shell is correct?
- It allows reading from many types of data sources
- All of the above (Correct answer)
- It runs/tests application code interactively
- It helps Spark applications to easily run on the command line of the system
Correct answer: All of the above
The Spark Shell is an interactive environment (REPL) that allows users to quickly experiment with Spark code, test applications, and perform interactive data analysis. It supports reading from various data sources and facilitates running Spark applications directly from the command line. Its versatility makes it an essential tool for development and exploration.
Question 25: What is the purpose of the external shuffle service in Spark?
- Maintains shuffle files on worker nodes so executors can be released during dynamic allocation (Correct answer)
- Streams shuffle data directly between executors bypassing the driver
- Compresses shuffle data to reduce network bandwidth usage
- Handles shuffle operations faster than in-executor processing
Correct answer: Maintains shuffle files on worker nodes so executors can be released during dynamic allocation
The external shuffle service runs on worker nodes and holds shuffle files, allowing executors to be released while preserving shuffle data.
Question 26: How do you access the underlying distributed data of a GraphX graph as standard RDDs?
- graph.toRDD()
- graph.repartition()
- graph.vertices and graph.edges (Correct answer)
- graph.toDataFrame()
Correct answer: graph.vertices and graph.edges
A GraphX Graph exposes its data directly via `graph.vertices` (VertexRDD) and `graph.edges` (EdgeRDD), which can be used as regular Spark RDDs.
Question 27: Which Spark MLlib algorithm performs dimensionality reduction?
- PCA (Principal Component Analysis) (Correct answer)
- GBTClassifier
- Naive Bayes
- KMeans
Correct answer: PCA (Principal Component Analysis)
PCA (Principal Component Analysis) reduces the dimensionality of feature vectors by projecting them onto principal components.
Question 28: The APIs for Apache Spark are
- Python
- All of the above (Correct answer)
- Java
- Scala
Correct answer: All of the above
Apache Spark provides high-level APIs in multiple programming languages to allow a broad range of developers to interact with its functionalities. These include Python (PySpark), Java, and Scala, making it versatile for different development environments. This broad support enhances its accessibility and adoption.
Question 29: What is the purpose of the Spark Web UI?
- Managing Spark application security and authentication
- Submitting new Spark jobs via a browser
- Monitoring running and completed jobs, stages, tasks, and executor metrics (Correct answer)
- Configuring Spark cluster resources
Correct answer: Monitoring running and completed jobs, stages, tasks, and executor metrics
The Spark Web UI (default port 4040) provides detailed metrics on jobs, stages, tasks, storage, and executors for debugging and optimization.
Question 30: Which method is used to run a SQL query on a registered temporary view in Spark SQL?
- spark.query()
- spark.sql() (Correct answer)
- spark.runSQL()
- spark.execute()
Correct answer: spark.sql()
spark.sql() executes a SQL query and returns the results as a DataFrame.
Databricks Certified Associate Developer for Apache Spark
The Databricks Certified Associate Developer for Apache Spark certification validates proficiency in using Apache Spark and the Databricks Lakehouse Platform to complete introductory-level data engineering tasks, covering Spark architecture, SQL, DataFrames, performance tuning, and advanced processing features.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds