← All Artificial Intelligence Flashcard Decks

Reinforcement Learning Flashcards

6 cards from real Artificial Intelligence practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 6 Reinforcement Learning flashcards as text
  1. What is Q-learning in reinforcement learning?

    Answer: A model-free algorithm that learns action-value (Q) functions to determine the best action in each state

    Q-learning learns Q(s,a) — the expected cumulative discounted reward of taking action a in state s — and derives the optimal greedy policy from these values.

  2. What does the discount factor (γ) control in reinforcement learning?

    Answer: The importance of future rewards relative to immediate rewards; values near 1 make the agent far-sighted

    A discount factor γ ∈ [0,1) exponentially reduces the weight of future rewards; γ near 0 makes the agent myopic (focus on immediate reward), while γ near 1 makes it long-sighted.

  3. What is the Bellman equation used for in reinforcement learning?

    Answer: Expressing the value of a state as the immediate reward plus the discounted value of the next state, forming the basis for dynamic programming

    The Bellman equation recursively decomposes the value function: V(s) = max_a [R(s,a) + γ Σ P(s'|s,a) V(s')], enabling iterative computation of optimal values.

  4. What is Deep Q-Network (DQN) and what was its breakthrough application?

    Answer: A deep RL algorithm that uses a neural network to approximate Q-values, first achieving human-level performance on Atari games

    DQN (DeepMind, 2015) uses a deep CNN to approximate Q(s,a) directly from raw pixel inputs, using experience replay and target networks to stabilize training on Atari 2600 games.

  5. What is 'experience replay' in DQN and why is it important?

    Answer: Storing past transitions in a buffer and sampling random mini-batches to break temporal correlations and stabilize training

    Experience replay stores (s, a, r, s') transitions in a replay buffer and samples random mini-batches, decorrelating consecutive training samples and improving data efficiency.

  6. What is the key idea behind policy gradient methods in reinforcement learning?

    Answer: Directly optimizing the parameters of a policy by computing gradients that increase the probability of high-reward actions

    Policy gradient methods parameterize the policy directly and update parameters by ascending the gradient of expected reward, computed via the log-probability of taken actions weighted by their returns.