Reinforcement Learning & Decision Making Flashcards
7 cards from real AML practice questions. Tap to flip, then mark Knew It or Still Learning — missed cards come back until you master them.
Read the first 7 Reinforcement Learning & Decision Making flashcards as text
What is the primary purpose of reward shaping in reinforcement learning?
Answer: To provide intermediate rewards that guide the agent more efficiently toward goals in sparse reward settings
Reward shaping adds supplementary intermediate rewards to the sparse environmental signal, providing denser feedback that helps agents learn faster without waiting for rare terminal rewards.
What is Inverse Reinforcement Learning (IRL) primarily designed to accomplish?
Answer: Inferring an unknown reward function from observed expert demonstrations
IRL infers the underlying reward function from expert behavioral demonstrations, which is useful when the reward function is unknown, hard to define, or too complex to specify manually.
What is the credit assignment problem in reinforcement learning?
Answer: Identifying which past actions were responsible for a reward received much later in a sequence
The credit assignment problem is the difficulty of determining which earlier actions in a long trajectory caused a delayed reward, requiring techniques like eligibility traces or TD(λ) to propagate credit backward.
In multi-agent reinforcement learning (MARL), why does non-stationarity pose a fundamental challenge?
Answer: Each agent's changing policy makes the environment appear non-stationary from any single agent's perspective
As each agent continuously updates its policy, the effective environment experienced by all other agents changes, violating the stationarity assumption that standard single-agent RL convergence proofs rely on.
What is the key difference between Monte Carlo and Temporal Difference (TD) methods for value estimation?
Answer: Monte Carlo methods update values only after a complete episode; TD methods update after each step using bootstrapped estimates
Monte Carlo methods compute actual returns from complete episode trajectories (unbiased but high variance), while TD methods bootstrap from current value estimates at each step (lower variance but biased).
What is the function of eligibility traces in reinforcement learning?
Answer: Maintaining a decaying record of recently visited state-action pairs to propagate credit backward over multiple steps
Eligibility traces keep a decaying record of recently encountered state-action pairs, allowing TD updates to assign credit to earlier states in proportion to their recency, bridging TD(0) and Monte Carlo as TD(λ).
What is the primary advantage of hierarchical reinforcement learning (HRL) for complex tasks?
Answer: It decomposes complex long-horizon tasks into sub-goals, improving sample efficiency and enabling sub-behavior reuse
HRL breaks long-horizon tasks into hierarchies where high-level policies set sub-goals for lower-level policies, dramatically improving sample efficiency and allowing learned sub-behaviors to transfer across tasks.