โ† All NLP Flashcard Decks

Advanced Topics & Theory Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Advanced Topics & Theory flashcards as text
  1. In seq2seq models, what problem does the attention mechanism primarily address?

    Answer: The fixed-length bottleneck of the encoder's final hidden state

    Attention lets the decoder dynamically focus on different encoder hidden states instead of compressing all source information into one vector.

  2. What distinguishes autoregressive language models (e.g., GPT) from masked language models (e.g., BERT)?

    Answer: GPT predicts the next token left-to-right; BERT predicts masked tokens using full bidirectional context

    GPT is a causal (unidirectional) model suited for generation, while BERT sees the entire context around each masked token.

  3. What is 'nucleus sampling' (top-p sampling) in text generation?

    Answer: Sampling from the smallest set of tokens whose cumulative probability exceeds p

    Top-p sampling accumulates tokens from highest to lowest probability until the cumulative mass reaches p, then samples from that dynamic set.

  4. Which metric evaluates machine translation quality by comparing n-gram overlap with reference translations?

    Answer: BLEU

    BLEU (Bilingual Evaluation Understudy) measures modified n-gram precision between hypothesis and reference translations with a brevity penalty.

  5. What is the 'exposure bias' problem in sequence-to-sequence training?

    Answer: At test time the model sees its own predictions, but during training it always sees ground-truth tokens

    Exposure bias arises because teacher-forcing at train time hides prediction errors, making the model fragile to its own mistakes at inference.

  6. Which approach to low-resource NLP adds a small number of trainable parameters to a frozen pre-trained model?

    Answer: Adapter layers

    Adapter layers insert lightweight bottleneck modules between transformer layers and only those new parameters are trained, preserving the original weights.

  7. In coreference resolution, what is an 'antecedent'?

    Answer: The earlier mention that a pronoun or noun phrase refers back to

    An antecedent is the previously mentioned entity to which a subsequent referring expression (like a pronoun) resolves.