Master of Data Science Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Master of Data Science flashcards as text
In Apache Spark, what is the key difference between a transformation and an action?
Answer: Transformations are lazy and return RDDs; actions trigger execution and return results
Spark transformations (like map, filter) are lazy and only define a computation graph, while actions (like collect, count) trigger actual execution.
Which technique is used to handle missing data by filling in plausible values based on observed data relationships?
Answer: Multiple imputation
Multiple imputation replaces missing values with several plausible estimates drawn from the observed data distribution, preserving statistical relationships.
What is the primary purpose of the attention mechanism in transformer models?
Answer: To allow the model to weigh the importance of different tokens when encoding each token
The attention mechanism computes a weighted sum of value vectors, where weights reflect how relevant each position is to the current token being processed.
In hypothesis testing, what does a p-value of 0.03 indicate?
Answer: There is a 3% chance of observing results this extreme if the null hypothesis is true
The p-value is the probability of obtaining a test statistic at least as extreme as observed, assuming the null hypothesis is true.
Which data structure is most efficient for implementing a priority queue?
Answer: Binary heap
A binary heap supports O(log n) insert and extract-min/max operations, making it the standard choice for priority queue implementations.
What is the role of the discriminator in a Generative Adversarial Network (GAN)?
Answer: To classify whether a sample is real or generated
The discriminator is trained to distinguish real data samples from fake ones produced by the generator, providing the adversarial training signal.
In dimensionality reduction, what does PCA maximize?
Answer: Explained variance in the projected data
PCA finds orthogonal components (principal components) that successively capture the maximum variance in the data.