Skip to content
VibeFormer

MODULE 12

Machine Learning Foundations

The concepts every algorithm shares: risk minimisation, generalisation, the bias–variance trade-off, validation and metrics.

20 lessons~9h reading

  1. 01

    What Machine Learning Actually Is

    Beginner

    Learning from data versus explicit programming, the three paradigms, and an honest account of what ML cannot do.

    22 min
  2. 02

    Formulating a Learning Problem

    Beginner

    Input and output spaces, hypothesis classes, loss functions, and turning a vague goal into an objective.

    Assumes: What Machine Learning Actually Is

    26 min
  3. 03

    Empirical Risk Minimisation

    Advanced

    True risk versus empirical risk, why we optimise a proxy, and the approximation–estimation decomposition.

    Assumes: Formulating a Learning Problem

    28 min
  4. 04

    Generalisation, Overfitting and Underfitting

    Beginner

    Diagnosing capacity problems from learning curves, with the classic polynomial-fit demonstration.

    Assumes: Formulating a Learning Problem

    28 min
  5. 05

    The Bias–Variance Trade-off

    Intermediate

    Full algebraic decomposition of expected squared error into bias, variance and noise, with a simulation.

    Assumes: Generalisation, Overfitting and Underfitting

    32 min
  6. 06

    The No Free Lunch Theorem

    Advanced

    Why no learner dominates across all problems, and what that means for model selection in practice.

    Assumes: The Bias–Variance Trade-off

    20 min
  7. 07

    VC Dimension and PAC Learning

    Advanced

    Shattering, VC dimension, sample complexity bounds, and the theory behind how much data is enough.

    Assumes: Empirical Risk Minimisation · Probability Inequalities

    34 min
  8. 08

    Train, Validation and Test Splits

    Beginner

    The role of each split, why the test set must stay untouched, and stratification.

    Assumes: Generalisation, Overfitting and Underfitting

    22 min
  9. 09

    Cross-Validation

    Intermediate

    k-fold, stratified, leave-one-out and nested CV, with the bias–variance trade-off in choosing k.

    Assumes: Train, Validation and Test Splits

    30 min
  10. 10

    Hyperparameter Search

    Intermediate

    Grid, random and Bayesian optimisation, successive halving, and budgeting search honestly.

    Assumes: Cross-Validation · Bayesian Optimisation

    28 min
  11. 11

    Classification Metrics

    Beginner

    Confusion matrix, accuracy, precision, recall, F1, specificity and Cohen's kappa, all computed by hand.

    Assumes: Train, Validation and Test Splits

    30 min
  12. 12

    ROC and Precision–Recall Curves

    Intermediate

    Threshold sweeps, AUC interpretation, and why PR curves beat ROC under heavy imbalance.

    Assumes: Classification Metrics

    28 min
  13. 13

    Regression Metrics

    Beginner

    MSE, RMSE, MAE, MAPE, R² and adjusted R², and which to report for which audience.

    Assumes: Train, Validation and Test Splits

    24 min
  14. 14

    Probability Calibration

    Advanced

    Reliability diagrams, Brier score, Platt scaling and isotonic regression.

    Assumes: ROC and Precision–Recall Curves

    26 min
  15. 15

    Handling Class Imbalance

    Intermediate

    Resampling, SMOTE, class weights, threshold tuning, and choosing metrics that survive skew.

    Assumes: ROC and Precision–Recall Curves

    28 min
  16. 16

    Feature Engineering

    Intermediate

    Transformations, interactions, binning, domain features, and why this still outperforms model tinkering.

    Assumes: Encoding Categorical Features

    30 min
  17. 17

    Feature Selection

    Intermediate

    Filter, wrapper and embedded methods; mutual information, RFE and stability selection.

    Assumes: Feature Engineering

    28 min
  18. 18

    Regularisation

    Intermediate

    L1 and L2 penalties, elastic net, the constrained-optimisation view, and why L1 induces sparsity.

    Assumes: The Bias–Variance Trade-off · Lagrange Multipliers

    30 min
  19. 19

    The Curse of Dimensionality

    Intermediate

    Volume concentration, distance concentration, sample-density collapse, and its consequences for kNN and kernels.

    Assumes: The Bias–Variance Trade-off

    26 min
  20. 20

    Pipelines and Data Leakage

    Intermediate

    The many ways leakage sneaks in — scaling before splitting, target encoding, temporal leaks — and how pipelines prevent it.

    Assumes: Cross-Validation

    28 min