MODULE 18
Large Language Models
How modern LLMs are built, aligned, decoded, evaluated, served and turned into agents — with the mechanics, not the hype.
32 lessons~16h reading
- 0126 min
What an LLM Actually Is
BeginnerComing soonNext-token prediction, scale, emergent behaviour, and a clear-eyed account of capabilities and limits.
Assumes: Encoder, Decoder and Encoder–Decoder Families
- 0234 min
Decoder-Only Architecture in Detail
AdvancedComing soonA modern LLM block dissected: pre-norm, RMSNorm, SwiGLU, GQA and the residual stream.
Assumes: What an LLM Actually Is · Transformer Shapes: End-to-End Walkthrough
- 0330 min
Pretraining Objectives
AdvancedComing soonCausal language modelling, the loss over a batch, teacher forcing, and packing sequences efficiently.
Assumes: Decoder-Only Architecture in Detail
- 0430 min
Pretraining Data Pipelines
AdvancedComing soonSourcing, deduplication, quality filtering, decontamination and data mixtures.
Assumes: Pretraining Objectives
- 0532 min
Scaling Laws
AdvancedComing soonKaplan and Chinchilla power laws, compute-optimal token budgets, and worked allocation arithmetic.
Assumes: Pretraining Data Pipelines
- 0630 min
Compute and Token Economics
AdvancedComing soonFLOPs per token, the 6ND rule, GPU-hour estimation, and inference cost per million tokens.
Assumes: Scaling Laws
- 0730 min
RoPE, ALiBi and Positional Schemes
AdvancedComing soonRotary embeddings derived, relative position bias, and how context extrapolation is achieved.
Assumes: Positional Encoding
- 0826 min
Context Windows
IntermediateComing soonThe quadratic cost of attention, long-context techniques, and what actually degrades at long range.
Assumes: RoPE, ALiBi and Positional Schemes
- 0930 min
The KV Cache
AdvancedComing soonWhy generation is memory-bound, cache size arithmetic, prefill versus decode, and paged attention.
Assumes: Context Windows
- 1030 min
Efficient Attention: MQA, GQA and FlashAttention
AdvancedComing soonSharing key/value heads, IO-aware attention kernels, and sparse and linear attention variants.
Assumes: The KV Cache
- 1130 min
Mixture of Experts
AdvancedComing soonSparse routing, top-k gating, load balancing losses, and the active vs total parameter distinction.
Assumes: Efficient Attention: MQA, GQA and FlashAttention
- 1232 min
Decoding Strategies
IntermediateComing soonGreedy, beam, temperature, top-k, top-p and min-p, with the same logits decoded every way by hand.
Assumes: Pretraining Objectives
- 1326 min
Speculative and Parallel Decoding
AdvancedComing soonDraft-and-verify acceleration, acceptance rates, and medusa/lookahead variants.
Assumes: Decoding Strategies
- 1430 min
Prompt Engineering Fundamentals
BeginnerComing soonInstruction clarity, role and delimiter use, few-shot exemplar selection, and output format control.
Assumes: Decoding Strategies
- 1528 min
Chain-of-Thought Reasoning
IntermediateComing soonEliciting intermediate steps, zero-shot CoT, self-consistency voting, and where CoT fails.
Assumes: Prompt Engineering Fundamentals
- 1630 min
Advanced Prompting Patterns
AdvancedComing soonReAct, tree-of-thought, least-to-most, self-refine and program-aided prompting.
Assumes: Chain-of-Thought Reasoning
- 1728 min
In-Context Learning
AdvancedComing soonWhat happens mechanistically when a model learns from the prompt, and induction heads.
Assumes: Chain-of-Thought Reasoning
- 1830 min
Structured Output and Function Calling
IntermediateComing soonJSON schema enforcement, constrained decoding, grammars, and tool-call protocols.
Assumes: Prompt Engineering Fundamentals
- 1926 min
Supervised Fine-Tuning: The Alignment Pipeline
IntermediateComing soonWhere SFT sits between pretraining and preference optimisation, and what each stage contributes.
Assumes: Pretraining Objectives
- 2030 min
Reward Modelling
AdvancedComing soonLearning from pairwise preferences, the Bradley–Terry model, and reward hacking.
Assumes: Supervised Fine-Tuning: The Alignment Pipeline
- 2134 min
RLHF with PPO
AdvancedComing soonThe full RLHF loop, the KL penalty against the reference policy, and its practical instabilities.
Assumes: Reward Modelling · TRPO and PPO
- 2232 min
DPO and Direct Preference Optimisation
AdvancedComing soonDeriving DPO from the RLHF objective, plus IPO, KTO, ORPO and SimPO compared.
Assumes: RLHF with PPO
- 2324 min
Constitutional AI and RLAIF
AdvancedComing soonReplacing human labels with model-generated critiques against an explicit set of principles.
Assumes: DPO and Direct Preference Optimisation
- 2430 min
Reasoning Models and Test-Time Compute
AdvancedComing soonLong chain-of-thought training, RL on verifiable rewards, and trading inference compute for accuracy.
Assumes: DPO and Direct Preference Optimisation
- 2528 min
Hallucination
IntermediateComing soonWhy next-token prediction fabricates, calibration and uncertainty, abstention, and grounding strategies.
Assumes: Reasoning Models and Test-Time Compute
- 2630 min
LLM Evaluation and Benchmarks
IntermediateComing soonMMLU, GPQA, HumanEval and friends; contamination, saturation and why leaderboards mislead.
Assumes: Hallucination
- 2726 min
LLM-as-Judge
AdvancedComing soonModel-graded evaluation, position and verbosity bias, rubric design, and agreement with humans.
Assumes: LLM Evaluation and Benchmarks
- 2832 min
Safety and Prompt Injection
AdvancedComing soonJailbreaks, direct and indirect prompt injection, data exfiltration risks, and layered defences.
Assumes: Structured Output and Function Calling
- 2930 min
Serving and Inference Optimisation
AdvancedComing soonContinuous batching, throughput vs latency, tensor parallelism, vLLM and capacity planning.
Assumes: The KV Cache
- 3032 min
LLM Agents
AdvancedComing soonThe perceive–plan–act loop, tool use, memory, reflection, and honest failure modes.
Assumes: Advanced Prompting Patterns · Informed Search and A*
- 3126 min
Multi-Agent Systems
AdvancedComing soonRole specialisation, debate, orchestration topologies, and when multi-agent is worse than one good prompt.
Assumes: LLM Agents
- 3230 min
Multimodal LLMs
AdvancedComing soonVision encoders, projection into token space, interleaved training, and audio and video extensions.
Assumes: LLM Agents · Self-Supervised and Contrastive Learning