← All CAIC Flashcard Decks

AI Tools & Technologies Flashcards

7 cards from real CAIC practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 AI Tools & Technologies flashcards as text
  1. Which vector database is most commonly used to store and retrieve embeddings in Retrieval-Augmented Generation (RAG) pipelines?

    Answer: Pinecone

    Pinecone is a purpose-built vector database widely used in RAG pipelines to store and perform similarity searches on embeddings.

  2. What is the primary role of LangChain in AI application development?

    Answer: Orchestrating LLM calls, memory, and tool integrations into chains

    LangChain is a framework that orchestrates LLM calls, memory management, and external tool integrations into composable chains.

  3. A client wants to deploy an open-source LLM on their own servers to avoid data leaving their network. Which tool best supports this use case?

    Answer: Ollama

    Ollama allows users to run open-source LLMs locally or on private servers, keeping data fully on-premises.

  4. Which technique in prompt engineering explicitly instructs a model to reason step-by-step before producing its final answer?

    Answer: Chain-of-thought prompting

    Chain-of-thought prompting encourages the model to show intermediate reasoning steps, improving accuracy on complex tasks.

  5. What does the 'temperature' parameter control in LLM inference?

    Answer: The randomness or creativity of model outputs

    Temperature scales the probability distribution over next tokens; higher values produce more random, creative outputs while lower values produce more deterministic ones.

  6. Which AI observability tool is specifically designed to trace, evaluate, and debug LLM applications end-to-end?

    Answer: LangSmith

    LangSmith, built by the LangChain team, provides tracing, evaluation, and debugging specifically for LLM-powered applications.

  7. When using OpenAI's API, which parameter limits the total number of tokens in the model's response?

    Answer: max_tokens

    The max_tokens parameter caps the number of tokens the model can generate in a single response.