Natural Language Processing Flashcards
7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 Natural Language Processing flashcards as text
What does the term 'word sense disambiguation' (WSD) refer to in NLP?
Answer: Determining which meaning of a polysemous word is intended in context
WSD is the task of automatically identifying the correct sense of an ambiguous word (e.g., 'bank' as financial institution vs. river bank) given its context.
In information retrieval, what does TF-IDF measure and why is the IDF component important?
Answer: How often a term appears in a document weighted by how rare it is across all documents; IDF down-weights common terms
TF-IDF multiplies term frequency (importance within a document) by inverse document frequency (rarity across corpus), giving high scores to terms that are distinctive for a document.
Which decoding strategy for sequence generation introduces randomness by sampling from the top-k most probable tokens at each step?
Answer: Top-k sampling
Top-k sampling restricts the sampling distribution to the k highest-probability tokens at each step, balancing diversity and coherence in generated text.
What is the primary benefit of using subword tokenization over word-level tokenization for NLP models?
Answer: Subword tokenization handles morphological variants and rare words without requiring an open vocabulary
By splitting rare or unknown words into frequent subword units, methods like BPE and WordPiece eliminate out-of-vocabulary issues while keeping vocabulary size tractable.
In the context of large language model fine-tuning, what does RLHF (Reinforcement Learning from Human Feedback) accomplish?
Answer: It aligns model outputs with human preferences by training a reward model and optimizing with PPO
RLHF first trains a reward model on human preference comparisons, then fine-tunes the LLM using PPO to maximize expected reward while minimizing KL divergence from the SFT policy.
Which property of contextualized embeddings makes them superior to static embeddings for tasks like named entity recognition?
Answer: They produce different vector representations for the same word depending on surrounding context
Contextualized embeddings (e.g., from BERT or ELMo) capture polysemy and contextual nuance by producing token representations that reflect the entire surrounding sentence.
What is the purpose of the [CLS] token in BERT and how is it typically used for classification tasks?
Answer: Its final hidden state aggregates sentence-level information and is fed to a classification head
BERT prepends [CLS] to every input; after encoding, its final-layer representation is treated as the pooled sentence embedding and passed to a linear classifier for sentence-level tasks.