โ† All MS-DS Master of Data science Flashcard Decks

Natural Language Processing Flashcards

7 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Natural Language Processing flashcards as text
  1. In coreference resolution, what does it mean for two mentions to be 'coreferent'?

    Answer: They refer to the same entity or event in the world

    Coreference resolution links noun phrases (e.g., 'Barack Obama' and 'he') that refer to the same real-world entity, forming coreference chains across a document.

  2. What distinguishes nucleus sampling (top-p sampling) from top-k sampling in text generation?

    Answer: Nucleus sampling dynamically selects the smallest set of tokens whose cumulative probability exceeds threshold p, adapting to the distribution shape

    Unlike top-k which fixes the number of candidates, top-p sampling includes tokens until their cumulative probability reaches p, allowing a larger or smaller candidate set depending on distribution sharpness.

  3. In multi-task learning for NLP, what is the key assumption that justifies sharing parameters across tasks?

    Answer: Tasks share useful linguistic features in their input representations, enabling positive transfer

    Multi-task learning assumes that related NLP tasks (e.g., POS tagging and NER) rely on overlapping syntactic or semantic features, so shared lower layers transfer knowledge beneficially.

  4. Which of the following is a key limitation of the BLEU metric when evaluating natural language generation quality?

    Answer: BLEU ignores recall and does not correlate well with human judgments for fluency or meaning preservation

    BLEU is a precision-focused n-gram overlap metric that ignores synonymy, paraphrase, and sentence-level meaning, resulting in poor correlation with human quality judgments on many generation tasks.

  5. What is 'catastrophic forgetting' in the context of sequential fine-tuning of pre-trained language models?

    Answer: The model's performance on previously learned tasks degrades significantly when fine-tuned on a new task

    When a neural network is fine-tuned on a new task, the parameter updates can overwrite knowledge from prior tasks, a phenomenon called catastrophic forgetting or catastrophic interference.

  6. In the encoder-decoder architecture for neural machine translation, what is the function of the cross-attention mechanism in the decoder?

    Answer: It allows the decoder to selectively attend to relevant encoder hidden states when generating each target token

    Cross-attention in the decoder uses decoder query vectors against encoder key-value pairs, enabling the model to focus on the most relevant source tokens at each generation step.

  7. Which regularization technique specific to transformer training helps stabilize optimization by normalizing activations before each sublayer's computation?

    Answer: Pre-layer normalization (Pre-LN) applied before each sublayer

    Pre-LN applies layer normalization to the input of each sublayer (attention or FFN) before the residual addition, leading to more stable gradient flow compared to the original Post-LN Transformer.

Natural Language Processing Flashcards โ€” MS-DS Master of Data science Study Cards with Answers