notebooks

Deep Reinforcement Learning — Exercises

Exercise course from bandits and value functions through Monte Carlo, off-policy learning and policy gradients to control in high-dimensional state spaces.

Exercise course accompanying the deep reinforcement learning lecture: multi-armed bandits, state and action values, generalised policy iteration, Monte Carlo methods, off-policy learning, policy gradients, and control in high-dimensional state spaces.