Spark SQL and DataFrames Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 Spark SQL and DataFrames flashcards as text
What is a Dataset in Spark compared to a DataFrame?
Answer: A Dataset is a strongly typed, object-oriented API; a DataFrame is an alias for Dataset[Row]
Dataset is a strongly typed distributed collection available in Scala/Java, while DataFrame is Dataset[Row] with untyped column access.
Which Spark SQL function fills null values in a DataFrame with a specified value?
Answer: df.na.fill()
df.na.fill() replaces null values with a specified value and can handle different column types.
What does the explode() function do in Spark SQL?
Answer: Converts a column of arrays or maps into multiple rows
explode() takes a column containing an array or map and creates a new row for each element.
Which format is recommended for reading structured data in Spark SQL when schema inference is needed?
Answer: JSON
JSON allows Spark to automatically infer the schema during read, though Parquet is preferred for performance in production.
What is the function of withColumn() in Spark DataFrames?
Answer: Adds a new column or replaces an existing column with the given column expression
withColumn() returns a new DataFrame by adding or replacing a column using the provided expression.
What does the repartition() function do to a Spark DataFrame?
Answer: Increases or decreases the number of partitions by performing a full shuffle
repartition() performs a full shuffle to create the specified number of partitions, evenly distributing the data.