LLMs, RAG, RLHF and Reasoning
How a language model actually works, from tokens to decoding, then fine-tuning and RLHF, then retrieval-augmented generation and its failure modes. The chapter where your own paper is the content rather than a footnote.
0. Why this is the chapter that matters most
Three reasons. It is your actual domain. It is what the panel is likeliest to probe, because it is where your published work sits. And France's competition authority already runs retrieval-augmented generation over its own case database, so this is not speculative — it is a deployed enforcement tool nobody has audited.
1. Tokens, the unit everything is built from
A model does not see words. It sees tokens — pieces of text from a fixed vocabulary, usually somewhere between a character and a word. Tokenisation is the process of cutting text into them.
| Term | Definition |
|---|---|
| Vocabulary | The fixed set of tokens the model knows, typically 30,000 to 150,000 of them |
| Byte-pair encoding | The common method. Start from characters and repeatedly merge the most frequent adjacent pair until you have your vocabulary. Frequent words become one token; rare ones get split |
| Token ID | Each token is really just an integer index into the vocabulary |
| Rule of thumb | In English, roughly 4 characters or 0.75 words per token |
2. What the model does with those tokens
You do not need the mathematics. You need the four ideas, in order.
| Step | What happens | Plain version |
|---|---|---|
| Embedding | Each token ID is replaced by a list of numbers — a vector — learned during training | Every token becomes a point in a space where nearby points mean similar things |
| Positional encoding | Information about *where* in the sequence each token sits is added | Otherwise the model sees a bag of tokens and *A sued B* looks identical to *B sued A* |
| Attention | For each token, the model computes how much to weight every other token, then builds a new representation as a weighted blend of them | Each word gets to look at every other word and decide which ones matter to its own meaning |
| Feed-forward layers, stacked | The blended representations pass through ordinary layers, and the whole attention-plus-layers block repeats dozens of times | Each repetition builds a slightly more abstract reading of the text |
One transformer block
| Architecture | Reads | Used for |
|---|---|---|
| Encoder-only, BERT-style | The whole text at once, in both directions | Classification, extraction, labelling a corpus. Still the right and far cheaper tool for coding 30 agency reports |
| Decoder-only, GPT-style | Left to right, each token seeing only what came before | Generation. Every model you think of as a chatbot |
| Encoder-decoder | Encodes an input, then generates an output | Translation, summarisation |
3. Pretraining, and why it works at all
Pretraining is the self-supervised step from the machine learning chapter, at enormous scale: take text, hide the next token, and train the model to predict it. The answer was free because you deleted it yourself. Repeat across trillions of tokens.
| Term | Definition |
|---|---|
| Parameters | The adjustable numbers inside the model. A 7-billion-parameter model has 7 billion of them |
| Compute, FLOPs | Total arithmetic used in training. The AI Act's systemic-risk threshold is set at 10 to the 25th FLOPs |
| Scaling laws | Empirical finding that performance improves predictably as parameters, data and compute grow together. Chinchilla showed most early models were badly undertrained — they needed far more data for their size |
| Emergence | Capabilities appearing abruptly beyond a certain scale rather than improving smoothly. Contested in detail, and central to Schrepel's argument that emergence defeats prediction |
| Mixture of experts | Only a fraction of the network activates per token, so a model can be large in parameters but cheap per token. One of your three subject models, gpt-oss-20b, is this design |
4. Decoding — turning probabilities into text
The model gives a probability for every possible next token. Decoding is how you choose one. This is a setting, not a property of the model, and it changes the output substantially.
| Method | What it does | Effect |
|---|---|---|
| Greedy | Always take the single most probable token | Deterministic in principle, and often repetitive and flat |
| Beam search | Keep several candidate continuations alive and pick the best overall sequence | Better for translation. Tends to produce bland text for open generation |
| Temperature | Flattens or sharpens the distribution before sampling. Below 1 sharpens toward the likeliest; above 1 flattens toward variety; 0 is effectively greedy | The single most important knob. Your paper used 0.0 throughout to minimise variation |
| Top-k | Sample only from the k most probable tokens | Cuts off the long tail of nonsense |
| Top-p, nucleus | Sample from the smallest set of tokens whose probabilities sum to p | Adapts how many options to consider based on how confident the model is |
| Repetition penalty | Down-weight tokens already used | Stops loops |
| Term | Definition |
|---|---|
| Context window | The maximum number of tokens the model can attend to at once — prompt plus generated output. Anything beyond it is invisible |
| KV cache | Stored intermediate results so generating token 500 does not recompute tokens 1 to 499. Why the first token is slow and the rest are fast |
| Lost in the middle | Measured tendency to use information at the start and end of a long context better than material buried in the middle. Directly relevant to stuffing 20 retrieved documents into a prompt |
5. Adapting a model: fine-tuning, SFT, RLHF and DPO
A pretrained model predicts plausible continuations. It does not follow instructions, refuse harmful requests, or adopt a tone. Those are added afterwards, in stages.
| Stage | What it is | Data needed |
|---|---|---|
| Pretraining | Next-token prediction over a vast corpus | Trillions of tokens, no labels |
| Supervised fine-tuning (SFT) | Continue training on example prompt-and-ideal-response pairs, so the model learns the shape of being helpful | Thousands to hundreds of thousands of written demonstrations |
| Reward modelling | Humans rank several candidate responses. A separate model is trained to predict those rankings — it learns to score responses as a human would | Human preference comparisons |
| RLHF | Reinforcement learning from human feedback. The language model is optimised to score highly on the reward model, using reinforcement learning — usually PPO | The reward model, plus compute |
| DPO | Direct preference optimisation. Skips the separate reward model and the reinforcement learning loop, optimising directly on preference pairs. Simpler, more stable, now very widely used | The same preference pairs |
| Constitutional AI / RLAIF | Uses a written set of principles and model-generated critiques in place of much of the human labelling | Principles, plus a capable model |
| Efficient method | What it does |
|---|---|
| Full fine-tuning | Update every parameter. Expensive, and needs the whole model in memory |
| LoRA | Freeze the original model and train small added matrices alongside it. A tiny fraction of the parameters, and the result is a small file you can swap in |
| QLoRA | LoRA on a quantised model, so a large model fits on a single consumer GPU |
| Quantisation | Store parameters at lower numerical precision — 8-bit or 4-bit instead of 16. Smaller and faster, with some quality loss |
| Distillation | Train a small model to imitate a large one's outputs |
6. Prompting
| Technique | Definition |
|---|---|
| System prompt | Instructions placed before the conversation, setting role and constraints. Usually hidden from the user |
| Zero-shot | Just ask |
| Few-shot | Include several worked examples in the prompt, so the model infers the pattern |
| Chain of thought | Ask the model to reason step by step before answering. Often improves accuracy on multi-step problems |
| Self-consistency | Sample several reasoning paths and take the majority answer |
| Structured output | Constrain the response to a schema, so it can be parsed rather than interpreted |
7. RAG — retrieval-augmented generation
RAG means fetching relevant documents and placing them in the prompt, so the model answers from them rather than from memory. Two phases.
The RAG pipeline, both phases
| Index time, done once | What happens |
|---|---|
| Load | Collect the documents — every decision from 1977 onward |
| Chunk | Split each into passages small enough to retrieve usefully |
| Embed | Convert each chunk into a vector capturing its meaning |
| Store | Put the vectors in a database built for similarity search |
| Query time, every question | What happens |
|---|---|
| Embed the query | Turn the question into a vector in the same space |
| Retrieve | Find the k nearest chunks by vector similarity |
| Rerank, optional | Re-score the candidates with a slower, more accurate model that reads query and chunk together |
| Augment | Paste the surviving chunks into the prompt |
| Generate | The model answers using what it was given |
| Term | Definition |
|---|---|
| Embedding | A vector representing meaning, where similar texts sit close together |
| Cosine similarity | The standard closeness measure between two vectors |
| Vector database | Storage optimised for nearest-neighbour search |
| ANN search | Approximate nearest neighbour. Trades exactness for speed, which means retrieval itself is approximate |
| BM25 | The classical keyword ranking function. Still strong, and the baseline any semantic method should be compared against |
| Hybrid search | Combining keyword and semantic retrieval, which usually beats either alone |
| Cross-encoder reranker | Reads query and candidate together rather than separately. Much more accurate, far slower, so used on a shortlist |
| Chunking strategies | Fixed-size; recursive, respecting paragraph boundaries; semantic, splitting on topic shifts; or parent-document, retrieving a small chunk but supplying its larger parent |
8. Reasoning in RAG
Basic RAG retrieves once and answers. That fails whenever the question needs several steps, or when the right search terms are not in the question.
| Technique | What it does |
|---|---|
| Query rewriting | Rephrase the user's question into better search terms before retrieving |
| Multi-query | Generate several reformulations, retrieve for each, and pool the results |
| HyDE | Have the model write a hypothetical ideal answer, then retrieve documents similar to *that* rather than to the question. Answers resemble documents more than questions do |
| Step-back prompting | Ask a more general question first to establish the governing principle, then the specific one |
| Multi-hop | Retrieve, read, work out what is still missing, retrieve again. Necessary for chained questions |
| Self-RAG | The model decides for itself whether to retrieve, and critiques whether what came back actually supports its answer |
| Corrective RAG | Assess retrieval quality and fall back — to a different query or to the open web — when it is poor |
| GraphRAG | Retrieve over a knowledge graph of entities and relations rather than over loose passages |
| Agentic RAG | The model plans a sequence of searches and tool calls rather than following a fixed pipeline |
9. The failure modes — and this is your section
| Failure | Definition |
|---|---|
| Hallucination | Fluent, confident output that is false. Follows directly from the training objective: plausibility was rewarded, truth was not |
| Prompt injection, direct | A user instructs the model to ignore its instructions |
| Prompt injection, indirect | The malicious instruction arrives inside retrieved content, so the attacker never touches the prompt |
| RAG poisoning, corpus poisoning | Planting material in the retrieval corpus to steer outputs |
| Sycophancy | Shifting the stated answer toward what the user seems to believe |
| Unfaithful chain of thought | Stated reasoning that does not describe the actual computation |
| Lost in the middle | Material buried in a long context being effectively ignored |
| Context conflict | Retrieved documents contradicting each other, or contradicting the model's training |
| Groundedness failure | An answer not supported by the documents it was given, while appearing to be |
10. Evaluating these systems
| Measure | What it asks |
|---|---|
| Faithfulness, groundedness | Is every claim in the answer supported by the retrieved documents? |
| Answer relevance | Does the answer address the question asked? |
| Context precision | Of what was retrieved, how much was actually relevant? |
| Context recall | Of what was needed, how much was retrieved? |
| Citation accuracy | Do the cited sources exist, and do they say what is claimed? |
| Refusal appropriateness | Does it decline when it should, and only then? |
LLM-as-a-judge means using a model to grade outputs. It beats surface-overlap metrics on semantic tasks, and it carries known biases.
| Judge bias | What it is |
|---|---|
| Position bias | Preferring whichever candidate appears first |
| Verbosity bias | Preferring longer answers regardless of quality |
| Self-preference | Preferring text generated by itself or its own family |
| Grading instability | The same input scored differently across calls, even at temperature zero |
11. Vocabulary, including the ones used without definition
| Term | Meaning |
|---|---|
| Foundation model, frontier model | A large general-purpose model trained at scale, adapted to many tasks downstream |
| Open weights versus open source | Open weights means the parameters are downloadable. Open source properly also implies the training data and code, which almost no model offers. Schrepel and Pott built a metric precisely because the distinction is routinely blurred |
| Inference | Running a trained model. Distinguished from training |
| Guardrails | Checks around a model constraining what it accepts or emits. On your CV from Astroformer |
| Red teaming | Deliberately attacking your own system to find failures. Now a legal obligation under AI Act Article 55 for systemic-risk models |
| Alignment | Making a model behave as intended. Alignment tax is capability lost in the process |
| Jailbreak | A prompt defeating a model's safety training |
| Agent, agentic | A model that plans, calls tools, and acts over several steps rather than answering once |
| Tool use, function calling | Letting a model invoke external functions — a search, a calculator, a database query |
| Regulatory sandbox | A legal term, not a technical one. A supervised environment where a firm may test an innovation under regulator oversight with some requirements relaxed. AI Act Articles 57 to 59 require Member States to establish them; the Digital Omnibus pushed that deadline to December 2027 |
| Model card, datasheet | Standardised documentation of a model's intended use, performance and limitations. Voluntary practice, and the ancestor of AI Act Annex IV technical documentation |
| Benchmark | A fixed dataset and metric for comparing systems. Benchmark contamination is when test data leaked into training |
| Ablation | Removing one component to measure its contribution. Your middle condition is an ablation: the fake document with the instruction removed |
| Drift | Performance decaying as the world moves away from the training data. Legal corpora drift constantly, because law changes |
12. How this connects to ATLANTIS
| Strand | The content of it here |
|---|---|
| Accuracy, the data problem | France's retrieval system is a deployed enforcement tool whose answers depend entirely on what its corpus contains and how it was chunked. Corpus completeness and chunk integrity are accuracy questions with no current standard |
| Fairness | Article 296 requires reasons and Article 14 requires meaningful oversight. Your finding is that a system's self-account can be true, checkable and incomplete — which attacks both at once, empirically |
| Institutional arrangements | Who audits an agency's retrieval corpus? Nobody, currently. There is no obligation, no methodology, and after the Annex III gap possibly no legal hook either |
13. What the panel brings to this chapter
| Panel member | Why this chapter is their territory |
|---|---|
| Thibault Schrepel | Teaches The Law of Artificial Intelligence. Wrote Decoding the AI Act. Built the EC decisions knowledge graph and offers to help agencies build queryable systems from their own records — which is this pipeline. Co-authored on foundation model competition with Pentland at MIT and on measuring model openness with Pott. His podcast has an episode on two language models learning to fix prices in a repeated duopoly, where some agents learn to cover their tracks — which is your research domain arriving from the substantive side |
| Catalina Goanta | Measurement over unstructured text at scale, and lawful research access to data held by someone unwilling to share it. Both are your problems |
| Georgiana Mirza | Digital ecosystems and data spaces, where model access and data access shape market structure |
| Tijmen Wisman | A system whose reasoning cannot be inspected from outside is the SyRI problem. He has seen a court strike one down for exactly that |
14. What is unexplored, and six projects
15. Your CV, mapped onto this chapter
| What you have | Where it lands | Why it fits |
|---|---|---|
| The Mens Rea Evaluator — controlled ground truth, three conditions, clean control, six probes, enforced measurement invariants | Projects 1, 2 and 6. The single most transferable asset you own | You have built and debugged the instrument, and the failure mode you found is the one Article 14 assumes away |
| RAG pipelines in production at 28,000 monthly users, with guardrails and edge cases found in live logs | Projects 1, 3 and 4 | Chunking and retrieval failures are things you have debugged rather than theorised |
| Adversarial prompting and red teaming, professionally at Outlier and in a published harness | Projects 1 and 2 | Article 55 makes this a statutory obligation. No translation needed |
| Fine-tuning on open-weight models — TinyLlama, Hugging Face, Gemma, Llama 3, NVIDIA NIM | Projects 1 and 6 | Credibility on the GPAI chapter and on the openness debate Schrepel writes about |
| Evaluation methodology — majority-voted judges with agreement recorded, deterministic parsing where no model is needed, invariants enforced in code | Projects 2 and 5 | This is the discipline the whole field is missing, and you have it in print |
| The v1 post-mortem — traces fed to a grader, empty outputs scored, an omitted control arm that invalidated the headline | Every project here | It is your best evidence that you audit your own work, and the best answer to a time something broke |
| Multilingual RAG across six Indian languages | Projects 1 and 3 | A pan-EU corpus is 24 equally authentic languages, and cross-language divergence is where it breaks first |
| Legal training in evidence and procedure | Projects 1, 2 and 5 | Why you can see that Article 86 and Article 27 of Regulation 1/2003 ask different questions about the same file |
16. If you remember twelve things
- A language model is a conditional distribution over the next token. Plausibility was rewarded; truth was not. Every failure follows from that.
- Tokens are not words, and legal citations fragment into meaningless debris.
- Attention lets every token weigh every other, which is what represents a negated agreement rather than an agreement.
- Temperature is the main decoding knob, and temperature zero is not determinism through an API.
- Pretraining, then supervised fine-tuning, then preference optimisation by RLHF or DPO.
- RLHF optimises what rankers preferred, which is where sycophancy comes from — and why persona-shift probes are necessary.
- Prompt, fine-tune, or retrieve. Only retrieval can produce a citation, which is why a body under a duty to give reasons needs it.
- RAG is chunk, embed, store, then retrieve, rerank, generate. Retrieval is approximate at every stage.
- Chain of thought can be unfaithful — it is not an explanation, and feeding it to a grader is what invalidated your v1.
- Indirect injection and corpus poisoning reach the model through retrieved content, never touching the prompt.
- Your finding: the document channel alone sufficed in 9 of 15, the clean control at zero is what licensed the claim, and self-report failed at every level.
- The real result is partial attribution — a true, checkable, incomplete reason that survives testing and points away from the cause.