Natural Language Processing Fundamentals Flashcards
7 cards from real DSE practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Natural Language Processing Fundamentals flashcards as text
Which technique is used in word2vec to predict surrounding context words given a target word?
Answer: Skip-gram
Skip-gram predicts surrounding context words given a center/target word, while CBOW does the opposite.
What does the term 'perplexity' measure in the context of language models?
Answer: How well a probability model predicts a sample
Perplexity measures how well a language model predicts a held-out sample — lower perplexity indicates a better model.
In transformer architectures, what is the purpose of positional encoding?
Answer: To inject sequence order information since attention is permutation-invariant
Self-attention has no inherent notion of order, so positional encodings are added to embeddings to convey token position.
Which evaluation metric is most appropriate for machine translation tasks?
Answer: BLEU score
BLEU (Bilingual Evaluation Understudy) measures n-gram overlap between generated translations and reference translations.
What is the main advantage of subword tokenization methods like Byte-Pair Encoding (BPE) over word-level tokenization?
Answer: Handles out-of-vocabulary words more gracefully
BPE breaks rare and unknown words into subword units, reducing out-of-vocabulary issues while keeping common words intact.
In Named Entity Recognition (NER), what does the BIO tagging scheme stand for?
Answer: Beginning, Inside, Outside
BIO tagging marks the Beginning of an entity, tokens Inside an entity, and tokens Outside any entity.
Which of the following best describes the attention mechanism's computational complexity relative to sequence length n?
Answer: O(n²)
Standard self-attention computes pairwise interactions between all tokens, resulting in O(n²) time and memory complexity.