โ† All Apache Spark Flashcard Decks

Spark SQL and DataFrames Flashcards

6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 Spark SQL and DataFrames flashcards as text
  1. What is a Dataset in Spark compared to a DataFrame?

    Answer: A Dataset is a strongly typed, object-oriented API; a DataFrame is an alias for Dataset[Row]

    Dataset is a strongly typed distributed collection available in Scala/Java, while DataFrame is Dataset[Row] with untyped column access.

  2. Which Spark SQL function fills null values in a DataFrame with a specified value?

    Answer: df.na.fill()

    df.na.fill() replaces null values with a specified value and can handle different column types.

  3. What does the explode() function do in Spark SQL?

    Answer: Converts a column of arrays or maps into multiple rows

    explode() takes a column containing an array or map and creates a new row for each element.

  4. Which format is recommended for reading structured data in Spark SQL when schema inference is needed?

    Answer: JSON

    JSON allows Spark to automatically infer the schema during read, though Parquet is preferred for performance in production.

  5. What is the function of withColumn() in Spark DataFrames?

    Answer: Adds a new column or replaces an existing column with the given column expression

    withColumn() returns a new DataFrame by adding or replacing a column using the provided expression.

  6. What does the repartition() function do to a Spark DataFrame?

    Answer: Increases or decreases the number of partitions by performing a full shuffle

    repartition() performs a full shuffle to create the specified number of partitions, evenly distributing the data.