Skip to content
VibeFormer

MODULE 18

Large Language Models

How modern LLMs are built, aligned, decoded, evaluated, served and turned into agents — with the mechanics, not the hype.

32 lessons~16h reading

  1. 01

    What an LLM Actually Is

    BeginnerComing soon

    Next-token prediction, scale, emergent behaviour, and a clear-eyed account of capabilities and limits.

    Assumes: Encoder, Decoder and Encoder–Decoder Families

    26 min
  2. 02

    Decoder-Only Architecture in Detail

    AdvancedComing soon

    A modern LLM block dissected: pre-norm, RMSNorm, SwiGLU, GQA and the residual stream.

    Assumes: What an LLM Actually Is · Transformer Shapes: End-to-End Walkthrough

    34 min
  3. 03

    Pretraining Objectives

    AdvancedComing soon

    Causal language modelling, the loss over a batch, teacher forcing, and packing sequences efficiently.

    Assumes: Decoder-Only Architecture in Detail

    30 min
  4. 04

    Pretraining Data Pipelines

    AdvancedComing soon

    Sourcing, deduplication, quality filtering, decontamination and data mixtures.

    Assumes: Pretraining Objectives

    30 min
  5. 05

    Scaling Laws

    AdvancedComing soon

    Kaplan and Chinchilla power laws, compute-optimal token budgets, and worked allocation arithmetic.

    Assumes: Pretraining Data Pipelines

    32 min
  6. 06

    Compute and Token Economics

    AdvancedComing soon

    FLOPs per token, the 6ND rule, GPU-hour estimation, and inference cost per million tokens.

    Assumes: Scaling Laws

    30 min
  7. 07

    RoPE, ALiBi and Positional Schemes

    AdvancedComing soon

    Rotary embeddings derived, relative position bias, and how context extrapolation is achieved.

    Assumes: Positional Encoding

    30 min
  8. 08

    Context Windows

    IntermediateComing soon

    The quadratic cost of attention, long-context techniques, and what actually degrades at long range.

    Assumes: RoPE, ALiBi and Positional Schemes

    26 min
  9. 09

    The KV Cache

    AdvancedComing soon

    Why generation is memory-bound, cache size arithmetic, prefill versus decode, and paged attention.

    Assumes: Context Windows

    30 min
  10. 10

    Efficient Attention: MQA, GQA and FlashAttention

    AdvancedComing soon

    Sharing key/value heads, IO-aware attention kernels, and sparse and linear attention variants.

    Assumes: The KV Cache

    30 min
  11. 11

    Mixture of Experts

    AdvancedComing soon

    Sparse routing, top-k gating, load balancing losses, and the active vs total parameter distinction.

    Assumes: Efficient Attention: MQA, GQA and FlashAttention

    30 min
  12. 12

    Decoding Strategies

    IntermediateComing soon

    Greedy, beam, temperature, top-k, top-p and min-p, with the same logits decoded every way by hand.

    Assumes: Pretraining Objectives

    32 min
  13. 13

    Speculative and Parallel Decoding

    AdvancedComing soon

    Draft-and-verify acceleration, acceptance rates, and medusa/lookahead variants.

    Assumes: Decoding Strategies

    26 min
  14. 14

    Prompt Engineering Fundamentals

    BeginnerComing soon

    Instruction clarity, role and delimiter use, few-shot exemplar selection, and output format control.

    Assumes: Decoding Strategies

    30 min
  15. 15

    Chain-of-Thought Reasoning

    IntermediateComing soon

    Eliciting intermediate steps, zero-shot CoT, self-consistency voting, and where CoT fails.

    Assumes: Prompt Engineering Fundamentals

    28 min
  16. 16

    Advanced Prompting Patterns

    AdvancedComing soon

    ReAct, tree-of-thought, least-to-most, self-refine and program-aided prompting.

    Assumes: Chain-of-Thought Reasoning

    30 min
  17. 17

    In-Context Learning

    AdvancedComing soon

    What happens mechanistically when a model learns from the prompt, and induction heads.

    Assumes: Chain-of-Thought Reasoning

    28 min
  18. 18

    Structured Output and Function Calling

    IntermediateComing soon

    JSON schema enforcement, constrained decoding, grammars, and tool-call protocols.

    Assumes: Prompt Engineering Fundamentals

    30 min
  19. 19

    Supervised Fine-Tuning: The Alignment Pipeline

    IntermediateComing soon

    Where SFT sits between pretraining and preference optimisation, and what each stage contributes.

    Assumes: Pretraining Objectives

    26 min
  20. 20

    Reward Modelling

    AdvancedComing soon

    Learning from pairwise preferences, the Bradley–Terry model, and reward hacking.

    Assumes: Supervised Fine-Tuning: The Alignment Pipeline

    30 min
  21. 21

    RLHF with PPO

    AdvancedComing soon

    The full RLHF loop, the KL penalty against the reference policy, and its practical instabilities.

    Assumes: Reward Modelling · TRPO and PPO

    34 min
  22. 22

    DPO and Direct Preference Optimisation

    AdvancedComing soon

    Deriving DPO from the RLHF objective, plus IPO, KTO, ORPO and SimPO compared.

    Assumes: RLHF with PPO

    32 min
  23. 23

    Constitutional AI and RLAIF

    AdvancedComing soon

    Replacing human labels with model-generated critiques against an explicit set of principles.

    Assumes: DPO and Direct Preference Optimisation

    24 min
  24. 24

    Reasoning Models and Test-Time Compute

    AdvancedComing soon

    Long chain-of-thought training, RL on verifiable rewards, and trading inference compute for accuracy.

    Assumes: DPO and Direct Preference Optimisation

    30 min
  25. 25

    Hallucination

    IntermediateComing soon

    Why next-token prediction fabricates, calibration and uncertainty, abstention, and grounding strategies.

    Assumes: Reasoning Models and Test-Time Compute

    28 min
  26. 26

    LLM Evaluation and Benchmarks

    IntermediateComing soon

    MMLU, GPQA, HumanEval and friends; contamination, saturation and why leaderboards mislead.

    Assumes: Hallucination

    30 min
  27. 27

    LLM-as-Judge

    AdvancedComing soon

    Model-graded evaluation, position and verbosity bias, rubric design, and agreement with humans.

    Assumes: LLM Evaluation and Benchmarks

    26 min
  28. 28

    Safety and Prompt Injection

    AdvancedComing soon

    Jailbreaks, direct and indirect prompt injection, data exfiltration risks, and layered defences.

    Assumes: Structured Output and Function Calling

    32 min
  29. 29

    Serving and Inference Optimisation

    AdvancedComing soon

    Continuous batching, throughput vs latency, tensor parallelism, vLLM and capacity planning.

    Assumes: The KV Cache

    30 min
  30. 30

    LLM Agents

    AdvancedComing soon

    The perceive–plan–act loop, tool use, memory, reflection, and honest failure modes.

    Assumes: Advanced Prompting Patterns · Informed Search and A*

    32 min
  31. 31

    Multi-Agent Systems

    AdvancedComing soon

    Role specialisation, debate, orchestration topologies, and when multi-agent is worse than one good prompt.

    Assumes: LLM Agents

    26 min
  32. 32

    Multimodal LLMs

    AdvancedComing soon

    Vision encoders, projection into token space, interleaved training, and audio and video extensions.

    Assumes: LLM Agents · Self-Supervised and Contrastive Learning

    30 min