Skip to content
VibeFormer

MODULE 02

Calculus and Optimisation

Single- and multi-variable calculus, convexity, Lagrange multipliers and KKT — the machinery behind every training loop.

17 lessons~7h reading

  1. 01

    Functions, Limits and Continuity

    BeginnerComing soon

    Limits from both sides, the epsilon–delta definition, continuity and the classification of discontinuities.

    20 min
  2. 02

    Differentiability and Rules of Differentiation

    BeginnerComing soon

    The derivative as a limit, differentiability vs continuity, and the product, quotient and chain rules.

    Assumes: Functions, Limits and Continuity

    22 min
  3. 03

    Rolle's and the Mean Value Theorems

    IntermediateComing soon

    The theorems that license almost every later bound, with the geometric picture and standard applications.

    Assumes: Differentiability and Rules of Differentiation

    18 min
  4. 04

    Taylor and Maclaurin Series

    IntermediateComing soon

    Polynomial approximation, the remainder term, and why second-order Taylor expansion underpins Newton's method.

    Assumes: Rolle's and the Mean Value Theorems

    26 min
  5. 05

    Maxima and Minima of One Variable

    BeginnerComing soon

    Critical points, first and second derivative tests, and distinguishing local from global extrema.

    Assumes: Differentiability and Rules of Differentiation

    22 min
  6. 06

    Optimisation in One Variable

    IntermediateComing soon

    Closed-form optimisation, boundary cases, and numerical line search methods including golden section.

    Assumes: Maxima and Minima of One Variable

    24 min
  7. 07

    Partial Derivatives and the Gradient

    IntermediateComing soon

    Partials, directional derivatives, the gradient as the direction of steepest ascent, and level sets.

    Assumes: Differentiability and Rules of Differentiation

    24 min
  8. 08

    The Jacobian and the Hessian

    AdvancedComing soon

    First- and second-order derivative matrices for vector-valued functions, and what Hessian definiteness tells you.

    Assumes: Partial Derivatives and the Gradient · Quadratic Forms and Definiteness

    26 min
  9. 09

    The Multivariable Chain Rule

    AdvancedComing soon

    Composing vector functions, the chain rule in matrix form, and its direct role in backpropagation.

    Assumes: The Jacobian and the Hessian

    24 min
  10. 10

    Matrix Calculus for Machine Learning

    AdvancedComing soon

    Derivatives with respect to vectors and matrices, layout conventions, and a reference table of standard identities.

    Assumes: The Multivariable Chain Rule

    28 min
  11. 11

    Convex Sets and Convex Functions

    IntermediateComing soon

    Convexity tests, Jensen's inequality, and why convex problems have no bad local minima.

    Assumes: The Jacobian and the Hessian

    26 min
  12. 12

    Lagrange Multipliers

    AdvancedComing soon

    Equality-constrained optimisation, the geometric meaning of multipliers, and shadow prices.

    Assumes: Convex Sets and Convex Functions

    26 min
  13. 13

    KKT Conditions

    AdvancedComing soon

    Inequality constraints, complementary slackness, and the conditions that define the SVM dual.

    Assumes: Lagrange Multipliers

    28 min
  14. 14

    Gradient Descent

    IntermediateComing soon

    The update rule, step-size selection, convergence on convex objectives, and failure modes.

    Assumes: Partial Derivatives and the Gradient

    28 min
  15. 15

    Newton and Quasi-Newton Methods

    AdvancedComing soon

    Second-order optimisation, Newton's method, and why BFGS and L-BFGS approximate the Hessian instead.

    Assumes: Gradient Descent · Taylor and Maclaurin Series

    26 min
  16. 16

    Constrained Convex Optimisation

    AdvancedComing soon

    Standard form problems, Lagrangian duality, weak and strong duality, and the duality gap.

    Assumes: KKT Conditions

    30 min
  17. 17

    Numerical Differentiation and Autodiff

    AdvancedComing soon

    Finite differences and their error, forward vs reverse mode automatic differentiation, and gradient checking.

    Assumes: The Multivariable Chain Rule

    26 min