Skip to content
VibeFormer

MODULE 16

Reinforcement Learning

Bandits and MDPs through to PPO: the full progression from tabular dynamic programming to deep policy-gradient methods.

26 lessons~13h reading

  1. 01

    The Reinforcement Learning Problem

    BeginnerComing soon

    Agent, environment, reward and the exploration–exploitation dilemma, contrasted with supervised learning.

    24 min
  2. 02

    Multi-Armed Bandits

    IntermediateComing soon

    The simplest RL problem: action values, regret, and incremental mean updates.

    Assumes: The Reinforcement Learning Problem

    28 min
  3. 03

    Exploration Strategies

    AdvancedComing soon

    Epsilon-greedy, optimistic initialisation, UCB and Thompson sampling, compared on regret.

    Assumes: Multi-Armed Bandits

    30 min
  4. 04

    Markov Decision Processes

    IntermediateComing soon

    States, actions, transitions, rewards and the Markov property; the MDP as the formal RL object.

    Assumes: The Reinforcement Learning Problem · Markov Chains

    30 min
  5. 05

    Returns, Policies and Value Functions

    IntermediateComing soon

    Discounted return, stochastic policies, and the state- and action-value functions.

    Assumes: Markov Decision Processes

    28 min
  6. 06

    The Bellman Equations

    AdvancedComing soon

    Expectation and optimality equations derived from first principles, in scalar and matrix form.

    Assumes: Returns, Policies and Value Functions

    32 min
  7. 07

    Iterative Policy Evaluation

    IntermediateComing soon

    Computing v_pi by repeated sweeps, with a full numeric trace on a grid world.

    Assumes: The Bellman Equations

    28 min
  8. 08

    Policy Iteration

    IntermediateComing soon

    Alternating evaluation and greedy improvement, the policy improvement theorem, and convergence.

    Assumes: Iterative Policy Evaluation

    28 min
  9. 09

    Value Iteration

    IntermediateComing soon

    Folding improvement into the update, the contraction-mapping argument, and stopping criteria.

    Assumes: Policy Iteration

    28 min
  10. 10

    Monte Carlo Methods

    IntermediateComing soon

    Learning from complete episodes, first-visit vs every-visit, and off-policy importance sampling.

    Assumes: Value Iteration

    30 min
  11. 11

    Temporal Difference Learning

    AdvancedComing soon

    TD(0), bootstrapping, the TD error, and the bias–variance comparison with Monte Carlo.

    Assumes: Monte Carlo Methods

    30 min
  12. 12

    SARSA

    AdvancedComing soon

    On-policy TD control, the quintuple update, and the cliff-walking comparison with Q-learning.

    Assumes: Temporal Difference Learning

    26 min
  13. 13

    Q-Learning

    AdvancedComing soon

    Off-policy control, the max operator, convergence conditions, and a Q-table updated step by step.

    Assumes: SARSA

    32 min
  14. 14

    Expected SARSA and Double Q-Learning

    AdvancedComing soon

    Reducing update variance, maximisation bias, and the double estimator fix.

    Assumes: Q-Learning

    24 min
  15. 15

    n-Step Methods and Eligibility Traces

    AdvancedComing soon

    Interpolating between TD and Monte Carlo, the lambda-return, and TD(lambda) with traces.

    Assumes: Q-Learning

    30 min
  16. 16

    Value Function Approximation

    AdvancedComing soon

    Replacing tables with parameterised functions, the deadly triad, and semi-gradient methods.

    Assumes: n-Step Methods and Eligibility Traces

    30 min
  17. 17

    Deep Q-Networks

    AdvancedComing soon

    Neural Q-functions, experience replay, target networks, and the Atari result.

    Assumes: Value Function Approximation · Adaptive Optimisers: AdaGrad to AdamW

    32 min
  18. 18

    Rainbow: DQN Improvements

    AdvancedComing soon

    Double DQN, duelling architectures, prioritised replay, noisy nets and distributional RL.

    Assumes: Deep Q-Networks

    28 min
  19. 19

    The Policy Gradient Theorem

    AdvancedComing soon

    Parameterising policies directly, the log-derivative trick, and the full theorem derivation.

    Assumes: Value Function Approximation

    34 min
  20. 20

    REINFORCE

    AdvancedComing soon

    Monte Carlo policy gradient, baselines for variance reduction, and the full algorithm.

    Assumes: The Policy Gradient Theorem

    28 min
  21. 21

    Actor–Critic Methods

    AdvancedComing soon

    Combining learned value estimates with policy gradients, and the advantage function.

    Assumes: REINFORCE

    30 min
  22. 22

    A2C and A3C

    AdvancedComing soon

    Synchronous and asynchronous parallel actor-critics, and why parallelism stabilises training.

    Assumes: Actor–Critic Methods

    24 min
  23. 23

    TRPO and PPO

    AdvancedComing soon

    Trust regions, the surrogate objective, the clipped ratio, and why PPO became the RLHF default.

    Assumes: Actor–Critic Methods

    34 min
  24. 24

    DDPG, TD3 and SAC

    AdvancedComing soon

    Continuous action spaces, deterministic policy gradients, and maximum-entropy RL.

    Assumes: TRPO and PPO

    30 min
  25. 25

    Model-Based Reinforcement Learning

    AdvancedComing soon

    Learning environment models, Dyna, planning with learned models, and MCTS.

    Assumes: Value Iteration

    30 min
  26. 26

    Offline Reinforcement Learning

    AdvancedComing soon

    Learning from fixed datasets, distributional shift, conservative Q-learning and behaviour cloning.

    Assumes: TRPO and PPO

    28 min