Master of Data Science Natural Language Processing 1 Flashcards
6 cards from real MS-DS Master of Data science practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 6 Master of Data Science Natural Language Processing 1 flashcards as text
Which algorithm is commonly used to segment text into subword units, helping models handle out-of-vocabulary words?
Answer: Byte Pair Encoding (BPE)
Byte Pair Encoding iteratively merges the most frequent character pairs to build a subword vocabulary, allowing models to represent rare or unseen words as combinations of known subword tokens.
In the transformer architecture, what does the term 'positional encoding' accomplish?
Answer: It injects information about each token's position in the sequence
Because transformers process all tokens in parallel without recurrence, positional encodings (sinusoidal or learned vectors) are added to embeddings so the model can distinguish token order.
Which evaluation metric computes the geometric mean of precision and recall, and is widely used for NLP classification tasks?
Answer: F1 score
The F1 score is the harmonic mean of precision and recall, giving a single balanced measure especially useful when class distributions are uneven.
What distinguishes a bidirectional LSTM from a standard (unidirectional) LSTM?
Answer: It reads the input sequence in both forward and backward directions simultaneously
A bidirectional LSTM runs two separate LSTMs — one left-to-right and one right-to-left — then concatenates their hidden states, giving each token context from both past and future tokens.
In the context of topic modeling, what does Latent Dirichlet Allocation (LDA) assume about documents?
Answer: Documents are mixtures of topics, and topics are mixtures of words
LDA is a generative probabilistic model that treats each document as a mixture of latent topics and each topic as a probability distribution over vocabulary words, ignoring word order (bag-of-words assumption).
Which pre-training objective does BERT use to learn deep bidirectional representations?
Answer: Masked language modeling combined with next sentence prediction
BERT is pre-trained with two objectives: Masked LM (randomly masking tokens and predicting them from surrounding context) and Next Sentence Prediction (predicting whether two sentences are consecutive), enabling rich bidirectional context.