Language Models Flashcards
7 cards from real NLP practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Language Models flashcards as text
What is 'quantization' in the deployment of large language models?
Answer: Representing model weights with lower-precision numbers (e.g., INT8 instead of FP32) to reduce size and speed up inference
Quantization compresses model weights by using fewer bits per parameter, significantly reducing memory and improving inference speed with minimal accuracy loss.
What is 'knowledge distillation' in the context of language models?
Answer: Training a smaller 'student' model to mimic the output distribution of a larger 'teacher' model
Knowledge distillation trains a compact student model using the soft probability outputs (logits) of a larger teacher model as training targets, transferring its learned representations.
What is the 'hallucination' problem in large language models?
Answer: The model generates text that is fluent and confident but factually incorrect or fabricated
Hallucination refers to language models generating plausible-sounding but false information, a key reliability challenge since models optimize for fluency, not factual accuracy.
In language modeling, what is the difference between 'top-k sampling' and 'greedy decoding'?
Answer: Greedy decoding always selects the most probable token; top-k samples from the k most probable tokens
Greedy decoding always picks the single highest-probability token (deterministic), while top-k restricts sampling to the k most probable tokens, introducing controlled randomness.
What is a 'mixture of experts' (MoE) architecture in the context of large language models?
Answer: A model architecture where each input token is routed to only a subset of specialized sub-networks (experts) rather than all parameters
In MoE models, a learned gating network routes each token to a sparse subset of 'expert' feed-forward layers, enabling very large total parameter counts without proportionally increasing compute.
What does 'grounding' a language model mean in NLP system design?
Answer: Connecting the model's outputs to verifiable external knowledge or real-world data to reduce hallucination
Grounding connects model outputs to external knowledge sources (databases, documents, APIs) so responses are anchored in verifiable facts rather than solely the model's parametric memory.
What is 'tokenization' in language models, and why is byte-pair encoding (BPE) commonly used?
Answer: Tokenization splits text into subword units; BPE iteratively merges frequent character pairs to balance vocabulary size with coverage of rare words
BPE learns a vocabulary of subword units by iteratively merging the most frequent adjacent character pairs, enabling models to handle rare and out-of-vocabulary words without an explosion in vocabulary size.