RL Bible

RL Bible · Appendix D

Annotated Bibliography

Every paper in the book, grouped by topic, with why it matters.

Every work cited in this book, grouped by topic, each with one line on why it matters. Chapter numbers point to where the work is taught (bold where it is a primary subject). All links verified at time of writing; DOIs and arXiv identifiers are permanent even where publisher pages gate access.

Textbooks, Courses, and Monographs

  • Sutton & Barto, Reinforcement Learning: An Introduction, 2nd ed. (MIT Press, 2018)incompleteideas.net/book/the-book-2nd.html — the field's spine and this book's constant companion for Parts 0–I. (Chs. 1–13)
  • Lattimore & Szepesvári, Bandit Algorithms (Cambridge UP, 2020)banditalgs.com — the definitive rigorous treatment of Chapter 2's material. (Ch. 2)
  • Szepesvári, Algorithms for Reinforcement Learning (Morgan & Claypool, 2010)PDF — sixty dense pages from MDPs to approximation; the short rigorous bridge. (Chs. 3, 21)
  • Puterman, Markov Decision Processes (Wiley, 1994)doi.org/10.1002/9780470316887 — the mathematical reference for everything Chapter 3 waved at. (Chs. 3–4)
  • Bertsekas & Tsitsiklis, Neuro-Dynamic Programming (Athena, 1996)athenasc.com/ndpbook.html — asynchronous DP theory and the first rigorous DP-to-learning bridge. (Chs. 4, 9)
  • Bertsekas, Reinforcement Learning and Optimal Control (Athena, 2019)web.mit.edu/dimitrib/www/RLbook.html — the control-theoretic retelling; cleanest frame for AlphaZero-style lookahead. (Chs. 4, 8)
  • Agarwal, Jiang, Kakade & Sun, RL: Theory and Algorithmsrltheorybook.github.io — the standard graduate theory text ("AJKS"); free. (Ch. 21)
  • OpenAI, Spinning Up in Deep RL (2018)spinningup.openai.com — the practitioner's companion to Parts II–III. (Chs. 1, 11–13)
  • Courses: David Silver (UCL/DeepMind) — davidsilver.uk/teaching; Berkeley CS285 (Levine) — rail.eecs.berkeley.edu/deeprlcourse; Stanford CS234 (Brunskill) — web.stanford.edu/class/cs234. Pair with Parts 0–II, II–IV, and the theory thread respectively.

Foundations, Bandits, and Exploration Theory

Classical RL: TD, MC, Traces, Planning

Function Approximation and the Deadly Triad

Deep Value-Based RL

Policy Optimization and Actor-Critic

Continuous Control and Maximum Entropy

Model-Based RL

Exploration in Deep RL

Imitation and Inverse RL

Offline RL

RL as Sequence Modeling

Goals, Hierarchy, Meta, Multi-Agent

RL for Language Models

Theory

Robotics: Classical and Deep

World Models

Vision-Language-Action Models