Ultimate AI Engineer Flashcards
7 cards from real AI practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Ultimate AI Engineer flashcards as text
In LoRA (Low-Rank Adaptation) fine-tuning, what is decomposed into low-rank matrices?
Answer: The weight update matrices ΔW
LoRA freezes original weights and approximates ΔW = A·B where A and B are low-rank, drastically reducing trainable parameters.
Which of the following is an example of few-shot prompting?
Answer: Including 3 example input-output pairs in the prompt before the actual query
Few-shot prompting embeds a small number of demonstration examples directly in the prompt to guide the model's output format and style.
A production LLM application has a P99 latency of 8 seconds. Which optimization most directly addresses time-to-first-token?
Answer: Applying KV-cache to the prefill step
Caching the KV states during prefill avoids recomputing them for the prompt, directly reducing time-to-first-token latency.
What problem does RLHF (Reinforcement Learning from Human Feedback) primarily solve?
Answer: Aligning model outputs with human preferences and values
RLHF trains a reward model on human preference data and uses RL to fine-tune the LLM to maximize that reward, improving alignment.
Which evaluation approach detects when a model performs well on a benchmark due to memorizing test data rather than generalizing?
Answer: Contamination analysis checking if test data appears in pretraining
Contamination analysis searches the pretraining corpus for n-gram overlaps with benchmark test sets to detect data leakage.
In a multi-agent AI system, what is the role of an 'orchestrator' agent?
Answer: Decomposing tasks and routing sub-tasks to specialized agents
An orchestrator decomposes complex goals into sub-tasks and coordinates specialized sub-agents to execute them.
What does 'context window' limit in a large language model?
Answer: The total tokens (prompt + completion) the model can process in one forward pass
The context window defines the maximum sequence length the model can attend to in a single inference call, encompassing both input and output tokens.