โ† All AI Flashcard Decks

Ultimate AI Engineer Flashcards

7 cards from real AI practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Ultimate AI Engineer flashcards as text
  1. What is 'prompt injection' in the context of deployed LLM applications?

    Answer: Malicious input that overrides system instructions to hijack model behavior

    Prompt injection attacks embed adversarial instructions in user input that override the developer's system prompt, causing the model to follow attacker directives.

  2. Which evaluation framework is specifically designed to assess LLM applications end-to-end using LLM-as-a-judge scoring?

    Answer: RAGAS for RAG pipelines or LangSmith evals

    RAGAS and LangSmith provide LLM-as-judge metrics like faithfulness, answer relevancy, and context precision tailored to generative AI pipelines.

  3. In the context of AI engineering, what does 'guardrails' refer to?

    Answer: Input/output validation layers that enforce safety and policy constraints around LLM calls

    Guardrails are programmatic checks (classifiers, rules, or another LLM) applied to inputs and outputs to detect and block policy violations.

  4. What is the primary benefit of using structured output (e.g., JSON mode) when calling an LLM in a production pipeline?

    Answer: It guarantees the model's output conforms to a parseable schema, removing brittle regex parsing

    Structured output constrains the model's decoding to valid JSON (or another schema), making downstream parsing reliable and eliminating format-related failures.

  5. Which strategy best handles a user query that exceeds an LLM's context window in a document QA system?

    Answer: Use hierarchical summarization or a retrieval step to select only relevant chunks

    Hierarchical summarization or retrieval-based selection ensures the most relevant content fits in the window without losing critical information through naive truncation.

  6. When deploying an AI model with Kubernetes, what resource should be explicitly requested in the pod spec to enable GPU scheduling?

    Answer: nvidia.com/gpu resource limit in the container spec

    Kubernetes requires explicit resource requests like `nvidia.com/gpu: 1` in the container spec so the scheduler assigns GPU-equipped nodes and the NVIDIA device plugin allocates the device.

  7. What distinguishes a 'zero-shot' evaluation from a 'few-shot' evaluation when benchmarking LLMs?

    Answer: Zero-shot provides no examples in the prompt; few-shot includes demonstration examples

    In zero-shot evaluation the model receives only the task instruction with no examples, while few-shot evaluation prepends k demonstration input-output pairs to the prompt.