Prepare for the NLP - Natural Language Processing exam with our free practice test modules. Each quiz covers key topics to help you pass on your first try.
Try these questions from our free NLP practice tests. The correct answer and an explanation follow each question.
Why might a tokenizer produce different numbers of tokens for the same English word depending on whether it appears at the start of a sentence or mid-sentence in GPT-style BPE?
Answer: B. GPT-style BPE treats a leading space as part of the token, so 'dog' and ' dog' are different tokens
GPT tokenizers prepend a space to most words as a prefix (e.g., 'Ġdog'), distinguishing a word at the start of a sentence from a mid-sentence occurrence.
What is the 'Pointwise Mutual Information' (PMI) method used for in sentiment analysis?
Answer: B. Estimating semantic orientation of phrases by comparing co-occurrence with positive vs. negative seed words
PMI measures how much more often a phrase co-occurs with positive seed words versus negative ones to estimate its sentiment orientation.
What is a 'sentence embedding,' and how does it differ from averaging individual word embeddings?
Answer: A. A sentence embedding is a single vector representing the whole sentence's meaning, often capturing word order and composition that simple averaging ignores
Models like Sentence-BERT produce sentence embeddings via fine-tuned transformers that encode word order and inter-word relationships, unlike naive averaging which is order-invariant.
Which post-editing metric measures the minimum number of edit operations needed to correct an MT output into an acceptable translation?
Answer: C. TER (Translation Edit Rate)
TER counts the number of shifts, insertions, deletions, and substitutions needed to convert the MT hypothesis into a reference translation.