โ† All NLP Flashcard Decks

Advanced Topics & Theory Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Advanced Topics & Theory flashcards as text
  1. What is the primary purpose of the Transformer's multi-head attention mechanism?

    Answer: To allow the model to jointly attend to information from different representation subspaces

    Multi-head attention runs attention in parallel across multiple subspaces, letting the model capture different types of relationships simultaneously.

  2. Which decoding strategy samples from the top-k most probable next tokens at each step?

    Answer: Top-k sampling

    Top-k sampling restricts the sampling pool to the k highest-probability tokens, balancing diversity and coherence.

  3. In the context of NLP, what does 'perplexity' measure?

    Answer: How well a probability model predicts a sample

    Perplexity is the exponentiated average negative log-likelihood of a test set, measuring how uncertain the model is about each token.

  4. What is 'catastrophic forgetting' in continual learning for NLP models?

    Answer: The tendency of a neural network to lose previously learned information when trained on new tasks

    Catastrophic forgetting occurs when optimizing for a new task overwrites the weights that encoded prior knowledge.

  5. Which technique uses a smaller 'student' model to mimic the output distribution of a larger 'teacher' model?

    Answer: Knowledge distillation

    Knowledge distillation trains a compact student model to match the soft probability outputs of a pre-trained teacher, compressing knowledge without large accuracy loss.

  6. What is the role of the key-query-value (K-Q-V) structure in self-attention?

    Answer: Queries match against keys to produce attention weights, then weighted sum of values forms the output

    Each token's query attends over all keys; the resulting attention weights are applied to values to produce a context-aware representation.

  7. Which regularization technique randomly masks input tokens during pre-training, requiring the model to reconstruct them?

    Answer: Masked language modeling (MLM)

    MLM, used in BERT, masks 15% of tokens and trains the model to predict the originals, learning bidirectional context.