Agent
Environment
Policy
State
Action Space
Discount Factor (Gamma)
Horizon
Markov Decision Process (MDP)
Bellman Optimality Equation
Credit Assignment
Sparse Reward
Model-Free RL
Policy-Based RL
Value-Based RL
Episodic Task
Centralized Training Decentralized Execution
Cooperative MARL
Multi-Agent RL (MARL)
Curiosity-Driven Exploration
Thompson Sampling
Adversarial Bandit
Bandit Problem
Contextual Bandit
EXP3
Regret
Stochastic Bandit
Generalized Policy Iteration
Policy Evaluation
Policy Improvement
Policy Iteration
Every-Visit MC
First-Visit MC
Importance Sampling (RL)
Monte Carlo Control
Maximization Bias
SARSA
Deadly Triad
Feature Engineering for RL
Function Approximation in RL
Generalization in RL
Linear Function Approximation
Radial Basis Function (RL)
Deep Q-Network (DQN)
Double DQN
Dueling DQN
Experience Replay
Prioritized Experience Replay
Rainbow DQN
Advantage Function
Baseline (Policy Gradient)
Deterministic Policy Gradient (DPG)
Policy Gradient Methods
Policy Gradient Theorem
Advantage Actor-Critic (A2C)
Asynchronous Advantage Actor-Critic (A3C)
Generalized Advantage Estimation (GAE)
DDPG
Distributional RL
IMPALA
PPO
SAC (Soft Actor-Critic)
TD3
AlphaZero
Monte Carlo Tree Search (MCTS)
MuZero
Planning with Learned Models
World Model
Reward Engineering
Reward Hacking
Reward Misspecification
Behavioural Cloning
DAgger
GAIL (Generative Adversarial Imitation Learning)
Imitation Learning
Inverse Reinforcement Learning (IRL)
Hierarchical RL
Meta-Reinforcement Learning
Off-Policy Evaluation
Offline RL
RLHF (RL from Human Feedback)
Transfer Learning in RL
Atari Learning Environment
DeepMind Lab
Gymnasium
-
G
Gymnasium documentation Art…
-
G
Gymnasium environment API reference Docs
MuJoCo
-
M
MuJoCo documentation — Overview Docs
-
G
Gymnasium MuJoCo environments Art…
OpenAI Gym
RLlib
-
D
RLlib documentation Docs
-
D
RLlib algorithm guide Docs
Stable Baselines3
Unity ML-Agents
Autonomous Driving
Game AI (Chess, Go, Atari)
Portfolio Optimization
Recommendation Systems (RL)
-
+1 more
-
P
Monte Carlo Methods — CSCI-531: every-visit MC prediction contrasted with first-visit Art… 4 terms
-
-
S
OpenAI Spinning Up — Part 3: Intro to Policy Optimization (Baselines in Policy Gradients) Art… 3 terms
-
O
Reinforcement Learning — Monte Carlo Methods: every-visit vs first-visit estimators and convergence Art… 3 terms
-
S
Reward and Return: discounted return and why future rewards are discounted by γ — OpenAI Spinning Up Art… 3 terms
-
W
Sutton & Barto, Chapter 4: Dynamic Programming — slides defining Generalized Policy Iteration PDF 3 terms
-
H
The Reinforcement Learning Framework: the agent-environment loop — Hugging Face Deep RL Course Docs 3 terms
-
P
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger) Pap… 2 terms
-
-
-
-
-
-
I
Markov Decision Processes: finite-horizon and infinite-horizon value functions — MIT 6.390 notes Art… 2 terms
-
-
-
-
-
A
QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning Pap… 2 terms
-
A
Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems (Bubeck and Cesa-Bianchi) Pap… 2 terms
-
A
9.5.2 Value Iteration — Artificial Intelligence: Foundations of Computational Agents (2nd ed.) Art… 1 term
-
A
9.5.3 Policy Iteration — Artificial Intelligence: Foundations of Computational Agents (2nd ed.) Art… 1 term
-
A
A Contextual-Bandit Approach to Personalized News Article Recommendation (Li, Chu, Langford, Schapire - LinUCB) Pap… 1 term
-
A
A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem Pap… 1 term
-
-
-
-
A
A Survey of Temporal Credit Assignment in Deep Reinforcement Learning (Pignatelli et al., 2023) Pap… 1 term
-
-
-
-
-
-
-
-
E
Bellman equation — Wikipedia (principle of optimality and the Bellman optimality equation) Ref… 1 term
-
L
Bellman Optimality Equations for V* and Q* — Lil'Log, A (Long) Peek into Reinforcement Learning Art… 1 term
-
-
-
-
-
A
Deep Radial-Basis Value Functions for Continuous Control (Asadi, Parikh, Parr, Konidaris & Littman) Pap… 1 term
-
-
-
-
A
Deep Reinforcement Learning with Double Q-learning (van Hasselt et al. — the Double DQN paper) Pap… 1 term
-
-
-
-
-
-
-
-
-
P
Efficient Off-Policy Meta-Reinforcement Learning via Probabilistic Context Variables (PEARL) Pap… 1 term
-
-
-
-
H
Finite-time Analysis of the Multiarmed Bandit Problem (Auer, Cesa-Bianchi, Fischer, 2002) - the UCB1 paper PDF 1 term
-
-
-
-
-
-
-
-
-
-
A
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures Pap… 1 term
-
S
Key Concepts and Terminology: agents, environments and the interaction loop — OpenAI Spinning Up Art… 1 term
-
-
-
R
Lecture 7: Function Approximation — RL Theory (linear value function approximation with basis functions) Art… 1 term
-
P
Leveraging Procedural Generation to Benchmark Reinforcement Learning (Procgen Benchmark, ICML 2020) Pap… 1 term
-
-
A
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm (AlphaZero) Pap… 1 term
-
-
C
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy (Ziebart PhD thesis) PDF 1 term
-
A
Monte Carlo Method for Learning State-Value Functions — First-Visit Method (tutorial with Python) Art… 1 term
-
-
-
-
-
-
-
-
-
-
-
L
Optimal Value and Policy: V*, Q* and the optimal policy π* — Lil'Log, A (Long) Peek into RL Art… 1 term
-
-
S
Policies: deterministic, stochastic and parameterized — OpenAI Spinning Up, Key Concepts in RL Art… 1 term
-
-
-
-
-
-
A
Reinforcement Learning with Function Approximation: From Linear to Nonlinear (Long & Han) Pap… 1 term
-
-
T
Reinforcement Learning, Part 8: Feature State Construction (polynomial, Fourier, aggregation, coarse/tile coding, RBFs) Art… 1 term
-
-
A
Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents Pap… 1 term
-
-
-
-
-
A
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor Pap… 1 term
-
-
D
Specification gaming: the flip side of AI ingenuity — Google DeepMind (reward misspecification) Art… 1 term
-
-
-
-
-
-
S
States and Observations: full vs partial descriptions of the world — OpenAI Spinning Up Art… 1 term
-
S
Stochastic Multi-Armed Bandit Problem and Algorithms (Oxford, Algorithmic Foundations of Learning, Lecture 15) PDF 1 term
-
G
Temporal difference reinforcement learning — SARSA (on-policy) compared with Q-learning Art… 1 term
-
-
-
-
-
-
M
Tile Coding — RL Sketchpad: offset tilings, binary features and gradient Monte Carlo approximation Art… 1 term
-
-
-
-
-
-
-
No terms or links match your search.