Model Evaluation & Optimization Techniques Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Model Evaluation & Optimization Techniques flashcards as text
What is the fundamental difference between model calibration and model discrimination?
Answer: Calibration measures predicted probability accuracy; discrimination measures ranking ability
Calibration measures how closely predicted probabilities match true event frequencies, while discrimination measures how well a model separates positive from negative cases (e.g., AUC).
Which hyperparameter search strategy combines random search with intelligent sampling by building a density estimator over promising regions?
Answer: Tree-structured Parzen Estimator (TPE)
TPE, used in Hyperopt, models the distribution of good and bad configurations separately using Parzen density estimators, directing search toward promising hyperparameter regions.
What does the coefficient of variation of RMSE across cross-validation folds indicate about a model?
Answer: The stability and reliability of the model's performance estimate
High variation in RMSE across folds suggests the performance estimate is unstable, indicating sensitivity to data partitioning and less reliable generalization estimates.
In gradient boosting, what is the role of the shrinkage parameter (learning rate)?
Answer: It scales each new tree's contribution, reducing overfitting at the cost of requiring more trees
The shrinkage/learning rate multiplies each tree's output before adding it to the ensemble, making each step conservative and reducing overfitting while necessitating more iterations.
A model's precision-recall curve lies entirely below another model's curve. What can you definitively conclude?
Answer: The higher model dominates the lower one across all recall thresholds
If one PR curve strictly dominates another (lies above it at all recall levels), the dominant model has higher precision at every recall value, making it unambiguously better for that task.
What is the primary risk of using the same validation set repeatedly for hyperparameter tuning?
Answer: Validation set overfitting — the model indirectly 'sees' the validation data through tuning decisions
Repeatedly evaluating hyperparameters on the same validation set causes the selection to adapt to that specific data partition, leading to optimistically biased performance estimates.
Which early stopping criterion monitors validation loss improvement over a patience window and is most robust to noisy training dynamics?
Answer: Stop when validation loss has not improved by at least delta in the last P epochs
Patience-based early stopping with a minimum delta tolerates short-term noise in validation loss while detecting when genuine improvement has stalled over P consecutive epochs.