← All NLP Flashcard Decks

Word Embeddings Flashcards

7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Word Embeddings flashcards as text
  1. What does 'debiasing' a word embedding mean in practice?

    Answer: Projecting embeddings to remove a defined bias direction, e.g., a gender axis, while preserving other semantic properties

    Bolukbasi et al. proposed identifying the gender subspace via PCA and projecting neutral words onto the orthogonal complement to neutralize gender bias.

  2. How are word embeddings typically initialized before fine-tuning on a downstream task?

    Answer: Loaded from a pre-trained model (e.g., GloVe or word2vec) rather than random initialization

    Using pre-trained embeddings as initialization leverages general language knowledge and speeds convergence, especially on small datasets.

  3. Why is Principal Component Analysis (PCA) sometimes applied to word embeddings after training?

    Answer: To reduce embedding dimensionality, remove noise, and sometimes improve downstream task performance

    PCA can compress 300-dimensional embeddings into fewer dimensions while preserving most variance, reducing memory and sometimes denoising the representation.

  4. What is the 'hubness problem' in high-dimensional word embedding spaces?

    Answer: A small number of points become nearest neighbors of many other points, skewing nearest-neighbor retrieval

    In high-dimensional spaces, some vectors systematically appear in many k-NN lists ('hubs'), degrading retrieval quality in cross-lingual or similarity search tasks.

  5. In the word2vec hierarchical softmax approach, what data structure is used to speed up vocabulary-scale predictions?

    Answer: A Huffman binary tree where frequent words are placed at shallower nodes

    Hierarchical softmax encodes the vocabulary in a Huffman tree; each word's probability is computed as a product of binary decisions along the path to its leaf, reducing cost from O(V) to O(log V).

  6. What is a 'sentence embedding,' and how does it differ from averaging individual word embeddings?

    Answer: A sentence embedding is a single vector representing the whole sentence's meaning, often capturing word order and composition that simple averaging ignores

    Models like Sentence-BERT produce sentence embeddings via fine-tuned transformers that encode word order and inter-word relationships, unlike naive averaging which is order-invariant.

  7. Which of the following best describes 'sparse' versus 'dense' word representations?

    Answer: Sparse representations (e.g., one-hot, TF-IDF) have mostly zero values; dense embeddings (e.g., word2vec) are short vectors with all non-zero values encoding distributed meaning

    One-hot vectors are as long as the vocabulary (often 100k+) with a single 1, while dense embeddings pack semantic information into compact 50–300 dimensional vectors with no zero structure.