โ† All Apache Spark Flashcard Decks

MLlib and Machine Learning Flashcards

6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 6 MLlib and Machine Learning flashcards as text
  1. What is the primary high-level API for machine learning in Apache Spark?

    Answer: spark.ml (DataFrame-based API)

    spark.ml is the primary, DataFrame-based API for machine learning in Spark, while spark.mllib is the older RDD-based library.

  2. What is a Pipeline in Spark MLlib?

    Answer: A sequence of stages (Transformers and Estimators) that form an ML workflow

    A Pipeline chains multiple Transformers and Estimators into a single ML workflow, simplifying training and prediction.

  3. What is the difference between a Transformer and an Estimator in Spark ML?

    Answer: A Transformer applies a transformation to a DataFrame; an Estimator fits data to produce a Transformer

    An Estimator has a fit() method that trains on data and returns a Transformer, which has a transform() method.

  4. Which Spark MLlib class is used to convert categorical string labels to numeric indices?

    Answer: StringIndexer

    StringIndexer encodes a string column of labels to a column of label indices, ordered by label frequencies.

  5. Which Spark MLlib algorithm is used for collaborative filtering-based recommendations?

    Answer: ALS (Alternating Least Squares)

    ALS (Alternating Least Squares) is Spark's collaborative filtering algorithm for building recommendation systems.

  6. What is the purpose of the VectorAssembler in Spark MLlib?

    Answer: Combines multiple feature columns into a single vector column

    VectorAssembler merges multiple numeric columns into a single feature vector required by Spark ML algorithms.