← All AML Flashcard Decks

Reinforcement Learning & Decision Making Flashcards

7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.

Read the first 7 Reinforcement Learning & Decision Making flashcards as text
  1. In Q-learning, what does the Q-value Q(s, a) represent?

    Answer: The expected cumulative discounted reward when taking action a in state s and following the optimal policy thereafter

    Q(s, a) is the expected total discounted return from taking action a in state s and then acting optimally, combining immediate reward with future expected returns.

  2. What fundamentally distinguishes SARSA from Q-learning?

    Answer: SARSA is on-policy and updates using the actual next action taken; Q-learning is off-policy and updates using the greedy max-Q action

    SARSA (on-policy) updates Q-values using the action actually taken by the behavior policy, while Q-learning (off-policy) always bootstraps from the greedy maximum Q-value action.

  3. What is the primary motivation for using Deep Q-Networks (DQN) over tabular Q-learning?

    Answer: DQN uses neural networks to approximate Q-values, enabling generalization over large or continuous state spaces

    DQN approximates Q-values with neural networks, enabling generalization across similar states and making RL tractable in high-dimensional or continuous state spaces where a Q-table is infeasible.

  4. What problem does experience replay solve in Deep Q-Networks?

    Answer: It reduces correlation between consecutive training samples, stabilizing neural network learning

    Experience replay stores transitions in a buffer and randomly samples mini-batches, breaking the temporal correlation between consecutive samples that would otherwise cause unstable, divergent training.

  5. In the Actor-Critic architecture, what is the primary role of the Critic?

    Answer: The Critic estimates a value function to evaluate the Actor's actions and reduce gradient variance

    The Critic learns a value function (V or Q) to assess how good the Actor's chosen actions are, providing a baseline that reduces variance in policy gradient estimates.

  6. What is the key innovation that makes Proximal Policy Optimization (PPO) stable during training?

    Answer: PPO clips the policy update ratio to prevent excessively large policy changes that could destabilize learning

    PPO constrains updates by clipping the probability ratio between the new and old policy within a trust region, preventing destructively large policy updates while keeping training computationally efficient.

  7. What is the 'deadly triad' in reinforcement learning, and why is it problematic?

    Answer: Function approximation, bootstrapping, and off-policy training — their combination can cause divergence and instability

    The deadly triad describes how combining function approximation, bootstrapping (TD updates), and off-policy learning introduces compounding approximation errors that can cause Q-value estimates to diverge.