MLlib and Machine Learning Flashcards
6 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 6 MLlib and Machine Learning flashcards as text
What is GBTClassifier in Spark MLlib?
Answer: A Gradient Boosted Trees Classifier for binary classification
GBTClassifier implements gradient boosted trees for binary classification, which iteratively trains decision trees to correct previous errors.
Which Spark MLlib algorithm is used for clustering unlabeled data?
Answer: KMeans
KMeans is an unsupervised clustering algorithm in Spark MLlib that groups data into k clusters based on feature similarity.
What does TrainValidationSplit do in Spark ML?
Answer: Trains multiple models on a parameter grid using a single train/validation split and selects the best
TrainValidationSplit evaluates a parameter grid using one train/validation split, which is faster but less reliable than CrossValidator.
Which feature transformer in Spark ML converts text documents to TF-IDF feature vectors?
Answer: HashingTF + IDF
HashingTF converts text tokens to term frequency vectors, and IDF scales by inverse document frequency to produce TF-IDF features.
What is the purpose of the IndexToString transformer in Spark ML?
Answer: Converts numeric predictions back to original string labels
IndexToString reverses the effect of StringIndexer, converting numeric label predictions back to the original string labels.
Which Spark MLlib algorithm is suitable for large-scale linear regression?
Answer: LinearRegression with L-BFGS optimizer
Spark's LinearRegression uses L-BFGS or OLS (ordinary least squares) optimization and scales well to large distributed datasets.