Agent
Environment
Policy
State
Action Space
Discount Factor (Gamma)
Horizon
Markov Decision Process (MDP)
Bellman Optimality Equation
Credit Assignment
Sparse Reward
-
S
Reward and Return: discounted return and why future rewards are discounted by γ — OpenAI Spinning Up Art… 3 terms
-
H
The Reinforcement Learning Framework: the agent-environment loop — Hugging Face Deep RL Course Docs 3 terms
-
I
Markov Decision Processes: finite-horizon and infinite-horizon value functions — MIT 6.390 notes Art… 2 terms
-
A
A Survey of Temporal Credit Assignment in Deep Reinforcement Learning (Pignatelli et al., 2023) Pap… 1 term
-
E
Bellman equation — Wikipedia (principle of optimality and the Bellman optimality equation) Ref… 1 term
-
L
Bellman Optimality Equations for V* and Q* — Lil'Log, A (Long) Peek into Reinforcement Learning Art… 1 term
-
-
-
-
-
S
Key Concepts and Terminology: agents, environments and the interaction loop — OpenAI Spinning Up Art… 1 term
-
L
Optimal Value and Policy: V*, Q* and the optimal policy π* — Lil'Log, A (Long) Peek into RL Art… 1 term
-
S
Policies: deterministic, stochastic and parameterized — OpenAI Spinning Up, Key Concepts in RL Art… 1 term
-
-
D
Specification gaming: the flip side of AI ingenuity — Google DeepMind (reward misspecification) Art… 1 term
-
S
States and Observations: full vs partial descriptions of the world — OpenAI Spinning Up Art… 1 term
-
-
No terms or links match your search.