← All NLP Flashcard Decks

Language Models Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Language Models flashcards as text
  1. What is 'transfer learning' in the context of NLP language models?

    Answer: Pretraining a model on large general corpora then fine-tuning it on a downstream task

    Transfer learning in NLP involves pretraining a model on a large dataset (e.g., Common Crawl) then adapting those learned representations to specific tasks with less data.

  2. What distinguishes a 'causal' language model from a 'non-causal' (bidirectional) one during training?

    Answer: Causal models mask future tokens so each position only attends to past tokens; bidirectional models attend to all positions

    Causal (autoregressive) models use a triangular attention mask to prevent each token from attending to future positions, making them suitable for generation.

  3. In language model training, what is 'teacher forcing'?

    Answer: Feeding ground-truth tokens as inputs at each decoding step rather than the model's own previous predictions

    Teacher forcing provides the ground-truth previous token as input at each training step, stabilizing gradients compared to feeding the model's own potentially wrong predictions.

  4. What is 'catastrophic forgetting' in the context of fine-tuning a pretrained language model?

    Answer: The model loses general knowledge acquired during pretraining when fine-tuned aggressively on a narrow task

    Catastrophic forgetting occurs when fine-tuning on task-specific data overwrites the broad representations learned during pretraining, degrading performance on other tasks.

  5. What does the acronym RLHF stand for, and how is it used in language model alignment?

    Answer: Reinforcement Learning from Human Feedback — trains models to produce outputs preferred by human raters

    RLHF uses human preference rankings to train a reward model, which then guides policy optimization (via PPO or similar) to make the language model more helpful and aligned.

  6. What is the role of 'Layer Normalization' in a Transformer language model?

    Answer: It stabilizes training by normalizing activations across the feature dimension within each layer

    Layer normalization standardizes activations within each layer to zero mean and unit variance, stabilizing training and allowing higher learning rates.

  7. What is 'in-context learning' as exhibited by large language models like GPT-3?

    Answer: The model adapting its behavior based on examples provided in the prompt without any weight updates

    In-context learning allows a model to perform new tasks by conditioning on a few examples in the prompt at inference time, with no gradient updates to model weights.