← All NLP Flashcard Decks

Text Preprocessing Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Text Preprocessing flashcards as text
  1. What is the 'vocabulary explosion' problem in NLP and how does subword tokenization address it?

    Answer: Unbounded vocabulary from rare/novel words; subword units cap vocabulary at a fixed size

    Subword tokenization (BPE, WordPiece) splits rare words into known subword units, enabling a fixed vocabulary to handle any new word.

  2. Which of the following describes the difference between type and token in corpus linguistics?

    Answer: A token is each individual word occurrence; a type is a unique word form

    Tokens are all word occurrences in text (including repeats), while types are the distinct unique words; 'the cat sat on the mat' has 6 tokens but 5 types.

  3. In preprocessing pipeline design, why is it generally recommended to fit vectorizers only on training data?

    Answer: Fitting on test data causes data leakage by letting the model learn test set statistics

    Fitting on test data leaks information about the test distribution into the model, causing overly optimistic evaluation metrics.

  4. What preprocessing technique is specifically designed to handle the informal spelling variations in user-generated text (e.g., 'goooood', 'pleaseeeee')?

    Answer: Character repetition normalization

    Character repetition normalization reduces elongated words (e.g., 'goooood' → 'good') by collapsing repeated characters to a standard form.

  5. How does the Punkt sentence tokenizer in NLTK handle abbreviations like 'Dr.' and 'Mr.'?

    Answer: It learns abbreviations from training data and treats their periods as non-sentence-final

    Punkt is an unsupervised algorithm that learns abbreviation patterns from text, avoiding false sentence splits on titles and abbreviations.

  6. What is the role of a 'text normalization pipeline' in an NLP system?

    Answer: To apply a sequence of preprocessing steps transforming raw text into a consistent, clean representation

    A text normalization pipeline chains steps like lowercasing, tokenization, stop word removal, and stemming/lemmatization into a reproducible preprocessing workflow.

  7. Why might preserving case information be important when preprocessing text for a Named Entity Recognition (NER) task?

    Answer: Capitalization is a strong signal for identifying proper nouns and named entities

    In English, named entities like person names, places, and organizations are typically capitalized, making case a valuable feature for NER models.