Master of Data Science Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Master of Data Science flashcards as text
What is the purpose of the 'EXPLAIN' command in SQL?
Answer: To show the query execution plan chosen by the optimizer
EXPLAIN reveals the execution plan the database engine will use to run a query, helping identify bottlenecks like full table scans.
Which ensemble method trains multiple models sequentially, where each new model corrects the errors of the previous one?
Answer: Gradient boosting
Gradient boosting builds models sequentially, with each tree fitting the residual errors (negative gradient of the loss) of the ensemble so far.
In NLP, what does TF-IDF measure?
Answer: The importance of a term in a document relative to how common it is across all documents
TF-IDF weights terms by their frequency in a document (TF) multiplied by the inverse of how many documents contain them (IDF), boosting rare but relevant terms.
What is a key advantage of using a data lake over a traditional data warehouse?
Answer: Ability to store raw, unstructured, and semi-structured data at low cost without requiring a predefined schema
Data lakes store data in its native format on cheap object storage, deferring schema definition to query time (schema-on-read), enabling flexible use of diverse data types.
Which Python library provides the DataFrame abstraction and is the standard tool for tabular data manipulation in data science?
Answer: pandas
pandas provides the DataFrame and Series objects with powerful data manipulation, aggregation, and I/O capabilities for tabular data.
In causal inference, what is the primary purpose of a randomized controlled trial (RCT)?
Answer: To eliminate confounding by randomly assigning subjects to treatment and control groups
Random assignment ensures that both observed and unobserved confounders are balanced across groups, enabling valid causal claims about the treatment effect.
What does the silhouette score measure in the context of clustering?
Answer: How similar each point is to its own cluster compared to other clusters
The silhouette score ranges from -1 to 1 and measures cohesion (intra-cluster distance) versus separation (nearest-cluster distance) for each point.