Basic Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Basic flashcards as text
Which of the following best describes a 'data pipeline'?
Answer: A series of automated steps to move and transform data
A data pipeline automates the flow of data through ingestion, transformation, and loading steps.
What does the term 'null hypothesis' mean in statistical testing?
Answer: The assumption that there is no effect or difference
The null hypothesis assumes no effect or difference exists; statistical tests try to find evidence against it.
Which metric measures the proportion of actual positives correctly identified by a classifier?
Answer: Recall
Recall (sensitivity) is TP / (TP + FN), measuring how well the model finds all true positives.
What is the purpose of a train/test split in machine learning?
Answer: To evaluate model performance on unseen data
Splitting data ensures the model is evaluated on data it has never seen, giving an honest performance estimate.
A correlation coefficient of -0.9 indicates:
Answer: Strong negative relationship
A correlation of -0.9 is close to -1, indicating a strong negative linear relationship between two variables.
Which of the following is NOT a common data cleaning task?
Answer: Training a neural network
Training a neural network is a modeling step, not a data cleaning task like imputation or deduplication.
What does ETL stand for in data engineering?
Answer: Extract, Transform, Load
ETL stands for Extract, Transform, Load — the standard process of moving data from source systems to a data warehouse.