Word Embeddings Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Word Embeddings flashcards as text
Which property of word2vec embeddings allows the analogy 'king - man + woman ≈ queen' to work?
Answer: Linear relationship encoding of semantic roles
Word2vec encodes semantic relationships as linear offsets in the embedding space, enabling arithmetic analogies.
What is the primary difference between the Skip-gram and CBOW architectures in word2vec?
Answer: Skip-gram predicts context from a target word; CBOW predicts a target from context words
Skip-gram takes a center word and predicts surrounding context words, while CBOW averages context words to predict the center word.
What does 'negative sampling' accomplish in word2vec training?
Answer: Approximates the softmax by training on a small set of noise words alongside the target
Negative sampling makes training tractable by updating weights only for the target word and a small random sample of non-target words.
In GloVe (Global Vectors), what is the main training objective?
Answer: Factorize the global word co-occurrence matrix into low-rank embeddings
GloVe directly factorizes the log of the global co-occurrence count matrix to produce word vectors.
Which evaluation method tests word embeddings by checking whether cosine similarity rankings match human-rated similarity scores?
Answer: Intrinsic evaluation via word similarity benchmarks
Intrinsic evaluation compares embedding-derived similarity rankings against human-annotated datasets like WordSim-353.
Why do subword-based embeddings like fastText outperform word2vec on morphologically rich languages?
Answer: They represent words as sums of character n-gram vectors, handling unseen word forms
FastText decomposes each word into overlapping character n-grams, so it can construct embeddings for out-of-vocabulary words with shared morphemes.
What does 'embedding dimensionality' control in a word embedding model?
Answer: The size of the dense vector representing each word, balancing expressiveness and memory
Dimensionality determines how many features each word vector has; higher dimensions can capture more nuance but require more memory and data.