Reinforcement Learning Lecture 7: Policy Gradient
David Silver’s Lecture 7. Optimize the policy directly, without going through value. The likelihood-ratio trick and the score function, REINFORCE, baselines and advantage, actor-critic, natural policy gradient, and a summary that ties six faces into one.