back

RL Bible

A from-scratch guide to reinforcement learning — bandits, Bellman equations, deep RL, offline RL, world models, and vision-language-action policies for robots.

Part 0 — Foundations

  1. Chapter 1What Reinforcement Learning Is
  2. Chapter 2Bandits & the Exploration Problem
  3. Chapter 3Markov Decision Processes

Part I — Tabular Methods

  1. Chapter 4Dynamic Programming
  2. Chapter 5Monte Carlo Methods
  3. Chapter 6Temporal-Difference Learning
  4. Chapter 7n-step Bootstrapping & Eligibility Traces
  5. Chapter 8Planning & Learning with Tabular Models

Part II — Approximation & Deep RL

  1. Chapter 9Value-Function Approximation
  2. Chapter 10Deep Q-Networks & the Value-Based Family
  3. Chapter 11Policy Gradients
  4. Chapter 12Advanced Policy Optimization
  5. Chapter 13Continuous Control & Maximum-Entropy RL
  6. Chapter 14Model-Based RL

Part III — The Modern RL Toolbox

  1. Chapter 15Exploration in Deep RL
  2. Chapter 16Imitation Learning & Inverse RL
  3. Chapter 17Offline (Batch) RL
  4. Chapter 18RL as Sequence Modeling
  5. Chapter 19Goal-Conditioned, Hierarchical, Meta & Multi-Agent RL
  6. Chapter 20RL for Large Language Models
  7. Chapter 21Theory & Foundations

Part IV — RL for Robotics

  1. Chapter 22Classical & Deep RL for Robots
  2. Chapter 23World Models
  3. Chapter 24Vision-Language-Action Models
  4. Chapter 25Frontier, Synthesis & Open Problems

Appendices

  1. Appendix ANotation & Conventions
  2. Appendix BMath Refresher
  3. Appendix CEnvironments & Code Setup
  4. Appendix DAnnotated Bibliography