← All NLP Flashcard Decks

Word Embeddings Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Word Embeddings flashcards as text
  1. Which phenomenon explains why 'bank' has different meanings in 'river bank' versus 'savings bank,' and how do static embeddings handle it?

    Answer: Polysemy; static embeddings assign one fixed vector per word regardless of context

    Static embeddings like word2vec average all senses of a polysemous word into a single vector, losing context-specific meaning.

  2. What role does the projection layer play in a word2vec neural network?

    Answer: It maps one-hot input vectors to dense lower-dimensional embedding vectors

    The projection layer is a simple linear lookup that converts sparse one-hot word representations into dense embedding vectors without a non-linearity.

  3. What is 'subsampling of frequent words' in word2vec, and why is it used?

    Answer: Randomly discarding high-frequency words during training to reduce noise and speed training

    Very common words like 'the' carry little specific semantic signal, so discarding them probabilistically improves both speed and embedding quality.

  4. How does the window size hyperparameter affect the semantic properties of learned word embeddings?

    Answer: Larger windows capture broader topical similarity; smaller windows capture tighter syntactic similarity

    A small context window (e.g., 2) makes the model focus on immediately adjacent syntactic patterns, while a large window (e.g., 10) captures document-level topic associations.

  5. In the context of word embeddings, what is 'retrofitting'?

    Answer: Post-hoc adjustment of pre-trained embeddings using semantic lexicons like WordNet

    Retrofitting pulls embeddings of synonymous words closer together using graph-based constraints from a semantic lexicon, improving lexical similarity.

  6. Which metric is most commonly used to measure the distance between two word embedding vectors?

    Answer: Cosine similarity

    Cosine similarity measures the angle between vectors and is preferred because it is invariant to vector magnitude, capturing directional semantic similarity.

  7. What does it mean for word embeddings to capture 'syntactic' regularities?

    Answer: Vector offsets encode grammatical relationships such as singular-plural or verb tense

    Word2vec and GloVe embeddings encode syntactic patterns like run→ran and dog→dogs as consistent directional offsets in the embedding space.