Reinforcement Learning Lecture 10: RL in Games

David Silver’s Lecture 10. A case study. Best response and Nash equilibrium seen through game theory, minimax search with a binary-linear value function, self-play reinforcement learning, combining minimax with RL, and imperfect-information games like poker. One recipe running through Chinook, Deep Blue, Logistello, TD-Gammon, and Maven.

June 19, 2021 · 16 min · rick

Reinforcement Learning Lecture 9: Exploration and Exploitation

David Silver’s Lecture 9. Use what you know now (exploitation), or find out more (exploration)? This note ties that dilemma into five principles, then follows the multi-armed bandit through regret and lower bounds, UCB, Thompson sampling, and information-state search, and on to contextual bandits and the extension to MDPs.

May 15, 2021 · 19 min · rick

Reinforcement Learning Lecture 8: Integrating Learning and Planning

David Silver’s Lecture 8. Learn a model of the environment from experience, then plan with that model. Model learning as supervised learning, the AB example, Dyna and Dyna-Q+, Monte-Carlo Tree Search and Go, TD search and Dyna-2.

April 17, 2021 · 5 min · rick

Reinforcement Learning Lecture 7: Policy Gradient

David Silver’s Lecture 7. Optimize the policy directly, without going through value. The likelihood-ratio trick and the score function, REINFORCE, baselines and advantage, actor-critic, natural policy gradient, and a summary that ties six faces into one.

March 13, 2021 · 6 min · rick

Reinforcement Learning Lecture 6: Value Function Approximation

David Silver’s Lecture 6. When states are too many to write down in a table, approximate the value with a handful of parameters. The incremental methods that learn features by gradient descent, mountain car and the bootstrapping debate, Baird’s counterexample where off-policy diverges, and the batch methods and DQN that pool experience.

February 13, 2021 · 13 min · rick

Reinforcement Learning Lecture 5: Model-Free Control

David Silver’s Lecture 5. Finding the optimal policy without a model. Why action values rather than state values, the two doors and ε-greedy and GLIE, Sarsa and Sarsa(λ) on the windy gridworld, importance sampling and off-policy learning, Q-learning through cliff walking, and the correspondence between DP and TD.

January 16, 2021 · 14 min · rick

Reinforcement Learning Lecture 4: Model-Free Prediction

David Silver’s Lecture 4. Estimate the value of a policy from experience alone, with no model of the environment. Monte-Carlo in Blackjack, temporal-difference learning on the drive home, the unified view of bootstrapping and sampling, and TD(λ) with eligibility traces.

December 12, 2020 · 6 min · rick

Reinforcement Learning Lecture 3: Dynamic Programming

David Silver’s Lecture 3. When the MDP is fully known, solve the Bellman equations by iteration to compute the optimal policy. Policy evaluation, policy iteration, and value iteration through a small grid world example.

November 14, 2020 · 6 min · rick

Reinforcement Learning Lecture 2: Markov Decision Processes

David Silver’s Lecture 2. The Markov decision process is the container that holds the reinforcement learning problem, built up one layer at a time, from the Markov chain through the reward process to the decision process, alongside Silver’s ‘student’ example.

October 17, 2020 · 6 min · rick

Reinforcement Learning, Lecture 1: The RL Problem

David Silver’s lecture 1. What sets reinforcement learning apart from other kinds of learning, reward and state, and the three pieces that make up an agent (policy, value function, model).

September 19, 2020 · 12 min · rick