Skip to content
VibeFormer

MODULE 24

AI Safety and Alignment

Specification gaming, reward hacking, interpretability, evaluations, red-teaming and control — the technical failure modes of capable systems, and what is actually known about mitigating them.

19 lessons~9h reading

  1. 01

    What AI Safety Actually Means

    BeginnerComing soon

    Separating the distinct problems bundled under one label: misuse, accidents, structural risk and misalignment, and why they need different responses.

    Assumes: What an LLM Actually Is

    26 min
  2. 02

    The Specification Problem

    IntermediateComing soon

    Why writing down what you want is harder than it looks: proxies diverge from intent under optimisation pressure, and Goodhart's law is the general statement of it.

    Assumes: What AI Safety Actually Means · Formulating a Learning Problem

    28 min
  3. 03

    Reward Hacking and Specification Gaming

    IntermediateComing soon

    Documented cases where agents maximised the stated reward while defeating its purpose, and what distinguishes a bug from a genuine optimisation success.

    Assumes: The Specification Problem · The Reinforcement Learning Problem

    30 min
  4. 04

    Goodhart's Law in Machine Learning

    AdvancedComing soon

    The four mechanisms — regressional, extremal, causal and adversarial Goodhart — and how each shows up in metrics, benchmarks and reward models.

    Assumes: Reward Hacking and Specification Gaming · Classification Metrics

    28 min
  5. 05

    Distributional Shift and Robustness

    AdvancedComing soon

    Why models fail when deployment differs from training, the i.i.d. assumption as a safety property, and what robustness guarantees are achievable.

    Assumes: What AI Safety Actually Means · Monitoring and Drift Detection

    30 min
  6. 06

    Adversarial Examples and Attacks

    AdvancedComing soon

    Imperceptible perturbations that flip predictions, why they exist geometrically, and the arms race between attacks and defences.

    Assumes: Distributional Shift and Robustness · Adversarial Machine Learning as a Game

    30 min
  7. 07

    Data Poisoning and Backdoors

    AdvancedComing soon

    Attacks on the training pipeline rather than the input: poisoned corpora, trigger-activated backdoors, and why web-scale pretraining data is hard to defend.

    Assumes: Adversarial Examples and Attacks · Data Curation for Fine-Tuning

    28 min
  8. 08

    The Alignment Problem

    AdvancedComing soon

    Outer and inner alignment stated precisely, why a system can pursue the trained objective and still not pursue the intended one, and what mesa-optimisation means.

    Assumes: Goodhart's Law in Machine Learning · RLHF with PPO

    32 min
  9. 09

    What RLHF Does and Does Not Fix

    AdvancedComing soon

    An honest account of preference-based alignment: what it demonstrably achieves, and the failure modes it cannot address — sycophancy, annotator bias and reward-model overoptimisation.

    Assumes: The Alignment Problem · DPO and Direct Preference Optimisation

    30 min
  10. 10

    Scalable Oversight

    AdvancedComing soon

    How to supervise a system on tasks humans cannot evaluate directly: debate, recursive reward modelling, weak-to-strong generalisation and their open problems.

    Assumes: What RLHF Does and Does Not Fix · Markov Games and Multi-Agent Learning

    30 min
  11. 11

    Interpretability: Reading the Model

    AdvancedComing soon

    Why black-box explanations are insufficient for safety, and the distinction between post-hoc explanation and mechanistic understanding.

    Assumes: The Alignment Problem · Model Interpretability

    28 min
  12. 12

    Mechanistic Interpretability

    AdvancedComing soon

    Features, circuits and superposition; induction heads, sparse autoencoders and activation patching as tools for finding what a network computes.

    Assumes: Interpretability: Reading the Model · Multi-Head and Masked Attention

    34 min
  13. 13

    Evaluations and Red-Teaming

    AdvancedComing soon

    Designing evals that measure dangerous capability rather than benchmark skill: capability versus propensity, elicitation, and why contamination makes most numbers unreliable.

    Assumes: What RLHF Does and Does Not Fix · LLM Evaluation and Benchmarks

    32 min
  14. 14

    Jailbreaks and Prompt Injection

    AdvancedComing soon

    Why instruction-following and instruction-refusing conflict, direct versus indirect injection, and why no purely prompt-level defence has held.

    Assumes: Evaluations and Red-Teaming · Safety and Prompt Injection

    30 min
  15. 15

    Safety of Tool-Using Agents

    AdvancedComing soon

    What changes when a model can act: irreversible actions, permission scoping, sandboxing, human-in-the-loop design and failure containment.

    Assumes: Jailbreaks and Prompt Injection · LLM Agents

    30 min
  16. 16

    AI Control

    AdvancedComing soon

    Designing deployments that remain safe even if the model is misaligned: monitoring, trusted-untrusted decomposition, and control evaluations.

    Assumes: Safety of Tool-Using Agents · Scalable Oversight

    28 min
  17. 17

    Dangerous Capability Evaluation

    AdvancedComing soon

    The capability domains that trigger heightened caution — cyber, bio, persuasion, autonomous replication — and how frontier safety frameworks set thresholds.

    Assumes: Evaluations and Red-Teaming

    28 min
  18. 18

    Safety Cases and Assurance

    AdvancedComing soon

    Making a structured, falsifiable argument that a system is safe enough to deploy, borrowing from aviation and nuclear practice.

    Assumes: Dangerous Capability Evaluation · AI Control

    26 min
  19. 19

    Open Problems, Honestly Stated

    AdvancedComing soon

    What is genuinely unsolved, where the field disagrees, which claims are overstated in both directions, and how to read the literature critically.

    Assumes: Safety Cases and Assurance · Mechanistic Interpretability

    28 min