โ† All AI Flashcard Decks

Ultimate AI Engineer Flashcards

7 cards from real AI practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Ultimate AI Engineer flashcards as text
  1. Which technique reduces LLM hallucinations by grounding responses in retrieved documents?

    Answer: Retrieval-Augmented Generation (RAG)

    RAG retrieves relevant documents at inference time and conditions the model's output on that grounded context, reducing hallucinations.

  2. In transformer attention, what does the 'key' vector represent?

    Answer: What information the token offers for matching

    Keys represent what each token offers; queries ask what is needed, and the dot product between them scores relevance.

  3. Which Python library is the industry standard for distributed training of large neural networks across multiple GPUs?

    Answer: PyTorch + DeepSpeed/FSDP

    PyTorch with DeepSpeed or FSDP (Fully Sharded Data Parallel) is the standard for large-scale distributed training.

  4. What is the primary purpose of a vector database like Pinecone or Weaviate in an AI system?

    Answer: Enabling fast approximate nearest-neighbor search over embeddings

    Vector databases index high-dimensional embeddings and support ANN search, enabling semantic retrieval at scale.

  5. An AI model deployed to production shows increasing latency over time without code changes. What is the most likely cause?

    Answer: Memory leak in the inference server

    A gradual latency increase without code changes typically points to a memory leak accumulating in the serving process over time.

  6. Which metric best evaluates the quality of generated text when a reference answer exists?

    Answer: BLEU or ROUGE score

    BLEU and ROUGE measure n-gram overlap between generated and reference text, making them standard for generation quality evaluation.

  7. What is 'temperature' controlling in LLM sampling?

    Answer: The randomness of token selection by scaling logits

    Temperature divides the logits before softmax: lower values make the distribution sharper (more deterministic), higher values flatten it (more random).