SophiArch
AdvancedML402

Reinforcement Learning

Master RL from tabular methods to deep policy optimisation. Build working agents in PyTorch, understand when each algorithm class applies, and develop the judgment to audit AI-generated RL configurations before they reach production.

Lessons are AI-assisted and human-reviewed. Learn more.

Syllabus

RL Foundations

01
The RL Problem and Markov Decision ProcessesFree preview
35 min
02
Value Functions and the Bellman Equation
40 min
03
Dynamic Programming: Policy and Value Iteration
45 min
04
Monte Carlo and Temporal-Difference Learning
40 min

Deep RL

05
Q-Learning
40 min
06
Deep Q-Networks
55 min
07
Policy Gradient Methods
50 min
08
Actor-Critic Methods and PPO
55 min