Text Preprocessing Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Text Preprocessing flashcards as text
What preprocessing challenge arises specifically when handling Chinese or Japanese text that doesn't exist with English?
Answer: Word boundary segmentation since words are not space-delimited
Chinese and Japanese text lacks spaces between words, so word segmentation algorithms are required before any word-level processing.
What is the effect of applying log normalization to term frequency (TF) values?
Answer: It reduces the impact of very high frequency terms by compressing the scale
Log normalization (e.g., 1 + log(tf)) prevents very frequent terms from dominating by compressing large frequency differences.
In preprocessing a legal or medical corpus, why might domain-specific stop words be added beyond the standard list?
Answer: Common domain terms like 'whereas' or 'patient' appear frequently but carry little discriminative value
High-frequency domain terms that appear in nearly every document (like 'patient' in medical text) behave like stop words and can be added to a custom list.
What is a 'character n-gram' and why is it useful in text preprocessing?
Answer: A sequence of n consecutive characters, useful for handling misspellings and morphology
Character n-grams capture subword patterns, making models robust to typos, morphological variants, and out-of-vocabulary words.
Which text normalization step would convert '5 kilometers' and '5 km' to a consistent representation?
Answer: Unit normalization or entity normalization
Unit normalization maps different representations of the same measurement (km, kilometers, kilometre) to a canonical form.
When preprocessing text for a sentiment analysis model, removing which of the following would likely hurt performance the most?
Answer: Negation words like 'not', 'never', 'hardly'
Negation words fundamentally flip sentiment polarity (e.g., 'not good' = negative), so removing them severely degrades sentiment classification.
What is the primary purpose of sentence boundary detection (SBD) in text preprocessing?
Answer: To split a document into individual sentences for downstream processing
SBD (also called sentence segmentation) correctly identifies where one sentence ends and another begins, enabling sentence-level analysis.