← All NLP Flashcard Decks

Language Models Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Language Models flashcards as text
  1. What is the 'context window' of a transformer-based language model?

    Answer: The maximum number of input tokens the model can attend to during inference

    The context window defines the maximum sequence length the model can process as input, limiting how much prior text it can condition on.

  2. In the Transformer architecture, what is the role of positional encodings?

    Answer: They inject information about the position of each token since self-attention is order-agnostic

    Since self-attention treats input tokens as a set (no inherent order), positional encodings add token position information so the model can learn order-dependent patterns.

  3. What is 'nucleus sampling' (top-p sampling) in language model decoding?

    Answer: Sampling from the smallest set of tokens whose cumulative probability exceeds a threshold p

    Top-p sampling dynamically selects a minimal set of tokens whose probabilities sum to at least p, then samples from that set — balancing diversity and coherence.

  4. What does 'instruction fine-tuning' accomplish in large language models?

    Answer: It trains the model on diverse instruction-response pairs to improve its ability to follow user directions

    Instruction fine-tuning teaches a pretrained model to interpret and respond to natural language instructions across a wide range of tasks.

  5. Which of the following is a major limitation of n-gram language models compared to neural language models?

    Answer: N-gram models suffer from data sparsity and cannot generalize across similar words

    N-gram models treat words as discrete symbols and cannot generalize — 'dog ran' and 'puppy ran' are completely unrelated — while neural models use dense embeddings that capture similarity.

  6. What is 'beam search' in language model generation, and what is its main drawback?

    Answer: Beam search maintains multiple candidate sequences simultaneously; it tends to produce generic, repetitive outputs

    Beam search tracks the top-b sequences at each step by score, but the highest-scoring sequences often converge to safe, repetitive outputs rather than diverse ones.

  7. What is the primary function of the feed-forward sublayer within each Transformer block?

    Answer: To apply a position-wise nonlinear transformation independently to each token representation

    The feed-forward sublayer applies two linear transformations with a nonlinearity (e.g., ReLU/GELU) independently to each token, adding representational capacity beyond what attention provides.