← All Apache Spark Flashcard Decks

Mixed Deck — All Apache Spark Topics Flashcards

100 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 20 Mixed Deck — All Apache Spark Topics flashcards as text
  1. Which language is used to create Spark?

    Answer: Scala

    Apache Spark was originally written in Scala, a powerful functional and object-oriented programming language. While it offers APIs in other languages, Scala remains its native and often most performant language for development. This foundational choice contributes to Spark's efficiency.

  2. What is a sliding window in Spark Structured Streaming?

    Answer: A window that moves forward by a fixed step smaller than the window duration, causing overlaps

    A sliding window has a duration and a slide interval; if the slide is smaller than the window duration, windows overlap.

  3. How does Spark handle executor failures during a job?

    Answer: Spark re-schedules the failed tasks on other available executors using RDD lineage

    When an executor fails, Spark reschedules its tasks on other executors and recomputes any lost RDD partitions using lineage.

  4. Which source reads data from Apache Kafka in Spark Structured Streaming?

    Answer: spark.readStream.format('kafka')

    spark.readStream.format('kafka') reads streaming data from Kafka topics using the Kafka connector for Structured Streaming.

  5. What is the purpose of spark.default.parallelism?

    Answer: Sets the default number of RDD partitions for transformations that create new partitions

    spark.default.parallelism sets the default number of partitions for RDD operations like reduceByKey and join when no explicit partition count is given.

  6. What is the purpose of the IndexToString transformer in Spark ML?

    Answer: Converts numeric predictions back to original string labels

    IndexToString reverses the effect of StringIndexer, converting numeric label predictions back to the original string labels.

  7. What is Apache Spark GraphX primarily used for?

    Answer: Graph-parallel computation and graph analytics

    GraphX is Spark's API for graph-parallel computation, enabling graph creation, transformation, and execution of graph algorithms at scale.

  8. Which join type is supported in Spark Structured Streaming for stream-to-static dataset joins?

    Answer: Inner, left outer, right outer, and full outer joins

    Stream-static joins in Structured Streaming support inner, left outer, right outer, and full outer joins.

  9. Which of the following statements regarding DataFrame is correct?

    Answer: DataFrames provide a more user-friendly API than RDDs.

    DataFrames offer a higher-level, more structured, and user-friendly API compared to RDDs, resembling tables in a relational database. They provide schema information and allow for optimized execution plans, making data manipulation and querying more intuitive and efficient for many use cases. This improved usability is a key advantage over raw RDDs.

  10. In Spark Structured Streaming, what does event time refer to?

    Answer: The time when the record was generated at the source

    Event time is the time embedded in the data itself, representing when the event actually occurred at the source.

  11. Which Spark MLlib algorithm performs dimensionality reduction?

    Answer: PCA (Principal Component Analysis)

    PCA (Principal Component Analysis) reduces the dimensionality of feature vectors by projecting them onto principal components.

  12. What happens when Spark runs out of memory for a shuffle and cannot fit data in memory?

    Answer: Spark spills the excess data to disk

    When execution memory is exhausted during a shuffle, Spark spills data to disk, which is slower but allows the job to continue.

  13. What is the Tungsten execution engine in Spark?

    Answer: A low-level execution engine that uses off-heap memory and code generation for performance

    Tungsten is Spark's physical execution engine that uses off-heap memory management and whole-stage code generation to maximize performance.

  14. Which of the following is not a Spark Ecosystem component?

    Answer: Sqoop

    Sqoop is a tool for transferring data between Hadoop and relational databases, and it is not a core component of the Apache Spark ecosystem. MLlib (Machine Learning), GraphX (graph processing), and BlinkDB (approximate query engine) are all part of or closely integrated with Spark. Therefore, Sqoop stands out as an external tool.

  15. Which class in Spark ML is used for hyperparameter tuning via cross-validation?

    Answer: CrossValidator

    CrossValidator performs k-fold cross-validation across a parameter grid to select the best model hyperparameters.

  16. What is dynamic partition pruning in Spark 3.x?

    Answer: Filtering partitions of a fact table at runtime using values from a dimension table join

    Dynamic partition pruning pushes the result of a dimension table filter into the fact table scan, skipping irrelevant partitions at runtime.

  17. What is the default number of partitions when creating an RDD from a collection using sc.parallelize()?

    Answer: Depends on the SparkContext default parallelism

    By default, sc.parallelize() uses the SparkContext's default parallelism, which is typically the number of cores.

  18. What is the purpose of the groupBy() function in Spark DataFrames?

    Answer: Groups rows by specified columns for aggregation

    groupBy() groups DataFrame rows by one or more columns and is typically followed by an aggregation function like agg(), count(), or sum().

  19. What is the key advantage of the `EdgePartition2D` strategy in GraphX over `EdgePartition1D`?

    Answer: It reduces vertex replication by hashing both source and destination IDs when assigning edges to partitions

    EdgePartition2D hashes both the src and dst vertex IDs together to assign edges, distributing vertex ghost copies more evenly and reducing total replication compared to 1D.

  20. What is data skew in Apache Spark and why is it a problem?

    Answer: Uneven data distribution across partitions causing some tasks to take much longer than others

    Data skew occurs when partitions have unequal sizes, making some tasks significantly slower and creating bottlenecks in the job.