Ultimate AI Engineer Flashcards
7 cards from real AI practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Ultimate AI Engineer flashcards as text
Which serving optimization packs multiple independent requests into a single GPU forward pass without waiting for a full batch to fill?
Answer: Continuous batching (iteration-level scheduling)
Continuous batching inserts new requests mid-generation at the token level, maximizing GPU utilization without fixed batch wait times.
An AI engineer wants to reduce an LLM's memory footprint by 4x with minimal accuracy loss. Which technique is most appropriate?
Answer: INT4 weight quantization
INT4 quantization represents weights in 4 bits instead of 16, achieving approximately 4x memory reduction with small accuracy degradation.
What is the key advantage of Flash Attention over standard attention implementation?
Answer: It computes attention in tiles to avoid materializing the full N×N attention matrix in HBM
Flash Attention tiles the computation to keep intermediates in SRAM rather than writing the full attention matrix to GPU high-bandwidth memory, dramatically reducing memory I/O.
Which technique enables an AI system to use external tools like calculators or web search during inference?
Answer: Function calling / tool use via structured output parsing
Function calling allows the model to emit structured JSON specifying a tool and arguments, which the host application executes and returns results for.
In MLOps, what does 'model drift' refer to?
Answer: Degradation in model performance caused by changes in real-world data distribution
Model drift occurs when the statistical properties of production inputs diverge from the training distribution, causing accuracy to degrade over time.
Which safety technique trains the model to refuse harmful requests by generating refusals as positive examples?
Answer: Constitutional AI / RLAIF using AI-generated critiques
Constitutional AI uses a set of principles and AI-generated critiques/revisions to self-improve responses toward safer behavior without human labelers for every example.
A RAG pipeline returns irrelevant chunks. Which component should an engineer investigate first?
Answer: The embedding model and chunking strategy
Retrieval quality depends on how well the embedding model represents semantic meaning and whether chunk size/overlap preserves coherent context.