Feature Engineering and Selection Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Feature Engineering and Selection flashcards as text
What does 'feature importance' from a gradient boosting model (e.g., XGBoost) typically measure?
Answer: The average gain in model performance when a feature is used in a split across all trees
XGBoost's gain-based feature importance averages the improvement in the loss function brought by each feature across all splits where it is used.
Which of the following correctly describes 'entity embeddings' as a feature engineering technique?
Answer: Learning dense low-dimensional vector representations for categorical features via a neural network
Entity embeddings train a neural network to map each category to a dense vector, capturing semantic relationships that one-hot encoding cannot.
In geospatial feature engineering, what is the purpose of computing Haversine distance as a derived feature?
Answer: To compute the great-circle distance between two latitude/longitude points on a sphere
The Haversine formula calculates the shortest distance over the Earth's surface between two coordinate pairs, creating a meaningful numeric proximity feature.
What is 'Boruta' and how does it differ from standard feature importance thresholding?
Answer: A feature selection algorithm that compares each feature's importance against shadow (permuted) features to determine statistical significance
Boruta creates randomized copies of features (shadows), trains a Random Forest, and accepts only features that consistently outperform the best shadow feature.
When engineering date/time features, which decomposition is most useful for capturing weekly seasonality in a demand forecasting model?
Answer: Day of week extraction
Day of week captures the recurring weekly pattern in demand (e.g., weekday vs. weekend effects) that is typically the strongest seasonal cycle in retail data.
What is the purpose of 'clipping' (winsorization) as a feature preprocessing step?
Answer: To cap extreme values at specified percentile thresholds, reducing the influence of outliers without deleting rows
Winsorization sets values below the lower percentile to that lower bound and values above the upper percentile to that upper bound, limiting outlier influence while retaining all rows.
Which method is best suited for selecting features when the relationship between features and the target is highly non-linear and non-monotonic?
Answer: Mutual information-based filter
Mutual information captures any statistical dependency between a feature and the target, including non-linear and non-monotonic relationships, unlike correlation which only detects linear ones.