MLlib and Machine Learning Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 MLlib and Machine Learning flashcards as text
What is the primary high-level API for machine learning in Apache Spark?
Answer: spark.ml (DataFrame-based API)
spark.ml is the primary, DataFrame-based API for machine learning in Spark, while spark.mllib is the older RDD-based library.
What is a Pipeline in Spark MLlib?
Answer: A sequence of stages (Transformers and Estimators) that form an ML workflow
A Pipeline chains multiple Transformers and Estimators into a single ML workflow, simplifying training and prediction.
What is the difference between a Transformer and an Estimator in Spark ML?
Answer: A Transformer applies a transformation to a DataFrame; an Estimator fits data to produce a Transformer
An Estimator has a fit() method that trains on data and returns a Transformer, which has a transform() method.
Which Spark MLlib class is used to convert categorical string labels to numeric indices?
Answer: StringIndexer
StringIndexer encodes a string column of labels to a column of label indices, ordered by label frequencies.
Which Spark MLlib algorithm is used for collaborative filtering-based recommendations?
Answer: ALS (Alternating Least Squares)
ALS (Alternating Least Squares) is Spark's collaborative filtering algorithm for building recommendation systems.
What is the purpose of the VectorAssembler in Spark MLlib?
Answer: Combines multiple feature columns into a single vector column
VectorAssembler merges multiple numeric columns into a single feature vector required by Spark ML algorithms.