Feature Engineering and Selection Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Feature Engineering and Selection flashcards as text
What is the main advantage of using quantile transformation over standard normalization for highly skewed data?
Answer: It maps values to a uniform or normal distribution, making it robust to outliers
Quantile transformation maps data to a target distribution based on rank, making it robust to extreme outliers unlike mean/std-based scaling.
In feature engineering for NLP, what does TF-IDF stand for and what does it measure?
Answer: Term Frequency–Inverse Document Frequency; measures how important a word is to a document relative to a corpus
TF-IDF weights a term by how often it appears in a document (TF) discounted by how common it is across all documents (IDF), highlighting distinctive words.
When would you prefer forward feature selection over backward feature elimination?
Answer: When the number of features is large relative to samples, making fitting a full model impractical
Forward selection starts with no features and adds one at a time, avoiding the need to fit a model on the full high-dimensional feature set.
What is 'Weight of Evidence' (WoE) encoding primarily used for?
Answer: Encoding categorical features in binary classification, particularly in credit scoring models
WoE encodes each category as the log ratio of the proportion of events to non-events, making it well-suited for logistic regression in credit risk models.
What problem does 'feature hashing' (the hashing trick) solve in machine learning pipelines?
Answer: Handles high-cardinality or unknown categorical values with a fixed-size feature vector
Feature hashing maps categories to a fixed-size vector using a hash function, allowing the model to handle unseen categories and large vocabularies without storing a full mapping.
What is the 'information gain' criterion in filter-based feature selection?
Answer: The reduction in entropy of the target variable given the knowledge of a feature's values
Information gain measures how much knowing a feature reduces uncertainty (entropy) about the target class, commonly used in decision tree feature ranking.
Why is it problematic to impute missing values using the global mean before splitting data into train and test sets?
Answer: It introduces data leakage because the imputation uses information from the test set
Computing the mean on the full dataset before splitting means the training imputation is informed by test set values, violating the independence of the test set.